{"id":"d751a559-20db-4a6b-87d8-015d4fec9532","arxiv_id":"2606.23215","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"LISDD localizes model discrepancies to operating regimes, selects sparse symbolic missing terms via holdout, and certifies them with sample-split F-tests and FDR control.","lead":"The paper introduces LISDD, a framework that finds where a trusted physics model fails in specific operating regimes, identifies a simple symbolic correction for the missing mechanism, and uses statistical tests to confirm the finding is real. This could improve hybrid physics-data models by avoiding global fixes that bias trusted parameters, especially in applications like building energy simulation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Clean-regime detection may not be robust when discrepancy is not localized or regimes overlap","rationale":"The reader's weakest assumption is precisely the load-bearing precondition for unbiased parameter recovery and calibrated testing; the controlled-experiment results do not address its violation. No other internal inconsistency is visible from the given description.","tokens_in":1796,"tokens_out":308,"duration_ms":18627,"concrete_test":"Generate synthetic trajectories in which the planted discrepancy term is active in every regime but with amplitude scaled by a smooth sigmoid of the operating variable; rerun the full LISDD pipeline and report (a) size of the largest detected clean regime and (b) resulting physical-parameter bias. If bias exceeds 0.05 or clean-regime fraction drops below 30 % while the F-test still declares significance, the headline performance does not generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The method first detects a clean regime to obtain unbiased physical-parameter estimates, then uses those to flag discrepancies via residual-energy and run the sample-split F-test. This sequence assumes the data contain a sufficiently large, automatically identifiable subset where the known model holds exactly. In the controlled experiments this is planted by construction, but the procedure offers no guarantee that detection succeeds when the missing mechanism is weak, present everywhere at low amplitude, or when operating regimes are not sharply separable. If detection errs, the subsequent bias reduction (0.002 vs 0.43) and exact detection claims no longer follow.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces LISDD, a framework for localizing model discrepancies in hybrid physics-data models. It automatically detects clean regimes to obtain unbiased estimates of trusted physical parameters, flags discrepant operating regimes via a residual-energy statistic, selects a sparse symbolic correction term by exhaustive holdout over a candidate library, and certifies significance with a sample-split F-test. An FDR extension handles multiple regions with distinct missing mechanisms. In controlled experiments with planted discrepancies, the method reports physical-parameter bias of 0.002 (vs. 0.43 for baselines), localization F1 of 0.80 (vs. 0.44), probability-one recovery of the correct symbolic form, exact detection, and control of multi-region false-discovery rate while recovering all planted mechanisms. The target application is grey-box building-energy models.","tokens_in":1944,"tokens_out":457,"duration_ms":20576,"significance":"If the central claims hold, LISDD supplies a statistically calibrated, localized diagnostic that avoids spreading local errors into clean regimes and biasing trusted parameters, which is a practical advance for hybrid modeling where physics is reliable only in subsets of the operating space. The use of finite-sample exact tests and explicit FDR control for multiple discoveries is a methodological strength.","major_comments":[{"comment":"Abstract: the reported bias reduction (0.002 vs. 0.43) and exact detection claims rest on the existence of an automatically detectable clean regime large enough for unbiased physical-parameter estimation. The manuscript provides no analysis or experiments showing that this detection step remains reliable when the missing mechanism is weak, present at low amplitude everywhere, or when operating regimes are not sharply separable; failure of this step would invalidate the subsequent performance numbers.","section":"Abstract"},{"comment":"Abstract: the term-selection step performs exhaustive holdout over a candidate library before applying the sample-split F-test. It is unclear whether the F-test is adjusted for the preceding combinatorial search or whether the reported probability-one recovery accounts for the effective multiple-testing burden induced by library size; this directly affects the claim of calibrated, identifiable discovery.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and for highlighting two important points about the scope of our claims. We address each comment below and commit to revisions that strengthen the statistical justification and empirical characterization of LISDD.","responses":[{"response":"We agree that the reported performance numbers presuppose successful identification of a sufficiently large clean regime. The controlled experiments in the manuscript use planted discrepancies of moderate strength with clearly separable regimes; they do not systematically vary discrepancy amplitude or regime overlap. We will add a dedicated subsection (and corresponding figures) that (i) sweeps the amplitude of the missing term from 0.1× to 2× the nominal scale, (ii) introduces controlled regime overlap via smoothed transition functions, and (iii) reports the empirical probability that the residual-energy detector recovers a clean regime large enough for unbiased parameter estimation. These results will be presented alongside the existing bias and F1 metrics so that readers can see the operating envelope of the method.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the reported bias reduction (0.002 vs. 0.43) and exact detection claims rest on the existence of an automatically detectable clean regime large enough for unbiased physical-parameter estimation. The manuscript provides no analysis or experiments showing that this detection step remains reliable when the missing mechanism is weak, present at low amplitude everywhere, or when operating regimes are not sharply separable; failure of this step would invalidate the subsequent performance numbers."},{"response":"The sample-split F-test is exact conditional on the term that was selected by the holdout procedure; it does not incorporate an explicit correction for the size of the candidate library. The probability-one recovery reported in the experiments is therefore an empirical observation under the specific library sizes and signal strengths tested, not a guarantee that holds after accounting for the combinatorial search. We will revise the methods section to state this conditioning explicitly, add a short theoretical remark on the induced multiple-testing burden, and include an ablation that varies library cardinality while tracking the empirical false-positive rate of the subsequent F-test. If the ablation reveals inflation, we will also report results with a simple Bonferroni adjustment applied post-selection.","revision_made":"partial","referee_comment":"[Abstract] Abstract: the term-selection step performs exhaustive holdout over a candidate library before applying the sample-split F-test. It is unclear whether the F-test is adjusted for the preceding combinatorial search or whether the reported probability-one recovery accounts for the effective multiple-testing burden induced by library size; this directly affects the claim of calibrated, identifiable discovery."}],"tokens_in":1509,"tokens_out":547,"duration_ms":16656,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that this method finds local model errors and certifies the fix, but only after locating a clean regime to fit the base physics.\n\nWhat stands out as new is the full sequence: auto-detecting a clean regime, using residual energy to spot bad spots, searching a library exhaustively for the missing term, certifying with split-sample F-test, and handling multiple regions with FDR. The abstract positions this as distinct from global discrepancy methods.\n\nThe experiments show clear gains: bias down to 0.002, F1 up to 0.80, and perfect recovery of the symbolic form in the planted cases.\n\nThe soft spot is exactly what the stress-test flags. The clean regime has to be there and detectable. If the missing physics is everywhere at low level or regimes blend, the initial fit gets contaminated and the rest doesn't work as claimed. Since the tests use constructed data with obvious separation, that part needs more scrutiny.\n\nThis is for folks doing hybrid physics and data models in applied settings like building energy. Someone who wants a diagnostic that points to specific missing mechanisms with a significance test would get something out of it.\n\nIt should go to a referee because the ideas on localization and certification are worth a close look, even with the assumption to verify.\n\nI'd say send it for review.","headline":"LISDD gives a pipeline to localize and certify missing physics terms but rests on detecting a clean regime first.","tokens_in":2426,"tokens_out":340,"would_cite":false,"duration_ms":32020,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LISDD localizes where a trusted physics model fails, recovers the exact symbolic missing mechanism, and certifies the discovery with exact finite-sample tests.","keywords":["model discrepancy","sparse discovery","localization","hybrid physics models","symbolic regression","statistical testing","grey-box modeling","building energy models"],"falsifier":"Run a controlled simulation with one planted local discrepancy; if LISDD either fails to recover the planted symbolic form with high probability or produces physical-parameter bias above 0.01, the central performance claims are falsified.","tokens_in":2694,"feed_emoji":"🔍","tokens_out":786,"duration_ms":18194,"temperature":0.7,"pith_summary":"The paper presents LISDD as a way to diagnose hybrid models by first locating a clean regime where the known physics holds exactly and estimating parameters there without contamination. It then identifies discrepant operating regimes via a residual-energy statistic, tests candidate symbolic corrections from a library using exhaustive holdout, and confirms each selected term with a sample-split F-test. An extension controls the false-discovery rate when multiple regions each have their own missing mechanism. This matters for applications such as building-energy models because global discrepancy fits can bias the trusted physics parameters and spread local errors into clean data, whereas the localized approach keeps parameter estimates clean while delivering statistically certified symbolic explanations.","feed_headline":"Method spots where physics breaks and recovers the exact missing equation","feed_subtitle":"LISDD first fits trusted parameters on a clean regime, then uses holdout and F-tests to identify local symbolic discrepancies while controll","key_machinery":"The LISDD procedure of clean-regime parameter estimation followed by local residual flagging, library-based holdout selection, and sample-split F-test certification.","core_discovery":"LISDD fits the known physics on an automatically detected clean regime, flags discrepant regions with a calibrated residual-energy statistic, selects the local missing term by exhaustive holdout over a candidate library, and confirms significance with a sample-split F-test. A false-discovery-rate extension handles multiple discrepant regions with different missing mechanisms. In controlled experiments, LISDD keeps physical-parameter bias at 0.002 versus 0.43 for global-discrepancy and black-box baselines, raises localization F1 from 0.44 to 0.80, recovers the correct symbolic form with probability one, attains exact detection, and controls the multi-region false-discovery rate while recoveri","pith_inferences":["The separation of clean-regime estimation from local search could be applied to other grey-box domains such as fluid or climate modeling where regime-dependent failures occur.","If clean regimes cannot be detected automatically, the method would require pairing with external regime-classification tools before the discrepancy stage.","The requirement for sample splitting and holdout implies a minimum data volume per regime that may limit use on very sparse observational datasets.","Certified local discoveries could feed directly into automated model-update pipelines that replace or augment the original physics law only in the affected regime."],"forward_implications":["Physical parameters stay unbiased at 0.002 error even when discrepancies exist in other regimes.","Localization F1 rises from 0.44 to 0.80 relative to global and black-box baselines.","The correct symbolic form of each missing mechanism is recovered with probability one.","Exact detection is achieved while the multi-region false-discovery rate remains controlled.","Every planted mechanism is recovered when several discrepant regions are present."],"fun_headline_variants":["Localizes where physics fails and finds sparse missing terms","Identifies discrepant regimes in hybrid physics models","Sparse symbolic discovery of local model discrepancies","Detects local physics errors using clean regime and holdout tests","Flags model errors and confirms missing terms with sample-split tests"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"An automatically detectable clean regime exists in which the known physics model holds exactly, allowing unbiased estimation of physical parameters before discrepancy search begins.","fun_headline_variants_meta":{"raw":{"variants":["Localizes where physics fails and finds sparse missing terms","Identifies discrepant regimes in hybrid physics models","Sparse symbolic discovery of local model discrepancies","Detects local physics errors using clean regime and holdout tests","Flags model errors and confirms missing terms with sample-split tests"]},"model":"grok-4.3","cost_usd":0.004995,"raw_usage":{"total_tokens":2502,"prompt_tokens":793,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":49949500,"prompt_tokens_details":{"text_tokens":793,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1636,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":793,"tokens_out":73,"duration_ms":17144,"temperature":1.0,"reasoning_tokens":1636,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T06:04:47.997145+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Run a controlled simulation with one planted local discrepancy; if LISDD either fails to recover the planted symbolic form with high probability or produces physical-parameter bias above 0.01, the central performance claims are falsified.","supporting_citations":[],"review_version":1}