{"id":"36c8309e-80e8-45ab-9749-d153475601ef","arxiv_id":"2601.05427","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A new anytime-valid inference method constructs test supermartingales to monitor consistency with Nash, correlated, and coarse correlated equilibria in repeated games and target policies in stochastic games.","lead":"The paper introduces a sequential testing framework using e-values and test supermartingales to detect real-time deviations from equilibrium behavior in multi-agent systems. This enables online monitoring of strategic play without fixed sample sizes, with extensions to different equilibrium concepts and stochastic games.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark equilibrium (or target policy) must be known exactly in advance; method provides no procedure for estimating it from data.","rationale":"The reader's weakest_assumption directly identifies the prerequisite that makes the supermartingale construction valid. No other internal inconsistency or missing step is visible from the provided claims.","tokens_in":1752,"tokens_out":305,"duration_ms":29740,"concrete_test":"Extract the explicit definition of the betting function (or e-value increment) in the repeated normal-form game section; recompute the one-step conditional expectation under the null using only the known equilibrium and payoff matrix. If this expectation exceeds 1 for any equilibrium type, the supermartingale claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central construction defines a test supermartingale by betting against the payoff inequalities that hold at the benchmark equilibrium (Nash, CCE, or CE). Under the null that play is consistent with that equilibrium, the conditional expectation of the increment must be non-positive for the supermartingale property to hold and deliver the claimed anytime-valid p-values and finite-time detection guarantees. This property is derived directly from the known payoff matrix and the known equilibrium mixed strategies (or correlation device). If the benchmark must instead be learned from the same stream of play, the martingale property is lost and the finite-time bounds no longer apply. The abstract and strongest claim treat the benchmark as given; no alternative construction or robustness result is indicated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces a sequential testing framework for detecting strategic deviations from a benchmark in multi-agent systems. It uses the e-value framework to construct test supermartingales that accumulate evidence against the benchmark (equilibrium in normal-form games or target policy in stochastic games). The approach unifies Nash, correlated, and coarse correlated equilibria with finite-time guarantees, provides online monitoring, applies Benjamini-Hochberg procedures for FDR control in large games, and extends to dynamic settings.","tokens_in":1908,"tokens_out":387,"duration_ms":14885,"significance":"If the supermartingale construction and finite-time bounds hold under the stated assumptions, the work offers a statistically sound, anytime-valid method for online detection of deviations that unifies multiple equilibrium concepts. The extension to stochastic games and FDR control are potential strengths for practical multi-agent monitoring.","major_comments":[{"comment":"Abstract and strongest claim: the supermartingale is constructed by betting against payoff inequalities that hold at a known benchmark equilibrium (or target policy). The conditional expectation of the increment is non-positive only when the benchmark is known exactly in advance and the observed play is tested directly against its payoff conditions. No procedure for estimating the benchmark from the data stream is indicated, which would break the martingale property and invalidate the finite-time guarantees and anytime-valid p-values.","section":"Abstract"}],"minor_comments":[{"comment":"Ensure that all finite-time detection time analyses and unification claims are accompanied by explicit theorem statements and proof sketches in the main body, with clear statements of the known-benchmark assumption.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript's scope appears limited to the known-benchmark case; this should be stated explicitly in the introduction to avoid overclaiming generality. No issues with citation patterns or scope fit noted."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for identifying the key assumption underlying our supermartingale construction. We address the major comment below.","responses":[{"response":"We agree that the supermartingale property holds precisely when the benchmark is known exactly in advance. This is the setting of the paper: the abstract states that we monitor consistency with 'a benchmark of strategic behavior' and verify adherence to 'a specified target policy, such as a computed equilibrium.' The full manuscript consistently treats the equilibrium (Nash, CCE, or CE) or target policy as given ex ante. No estimation procedure from the data stream is proposed or claimed, precisely because it would invalidate the martingale property. Our contribution is limited to the known-benchmark case, where the finite-time guarantees and anytime-valid p-values are valid. We have added one clarifying sentence in the introduction to make this assumption even more explicit.","revision_made":"partial","referee_comment":"[Abstract] Abstract and strongest claim: the supermartingale is constructed by betting against payoff inequalities that hold at a known benchmark equilibrium (or target policy). The conditional expectation of the increment is non-positive only when the benchmark is known exactly in advance and the observed play is tested directly against its payoff conditions. No procedure for estimating the benchmark from the data stream is indicated, which would break the martingale property and invalidate the finite-time guarantees and anytime-valid p-values."}],"tokens_in":1291,"tokens_out":311,"duration_ms":17346,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main move is to take the e-value supermartingale construction for safe anytime inference and use it to monitor repeated play against a fixed equilibrium benchmark. By betting against the payoff inequalities that hold at the benchmark, they produce a process that grows when observed payoffs violate the equilibrium conditions, giving an interpretable online measure of departure with finite-time guarantees. The same template covers Nash, correlated, and coarse correlated equilibria by changing which inequalities serve as the null, and they carry the idea over to stochastic games by testing trajectories against a specified target policy. They also add a Benjamini-Hochberg layer for controlling false discoveries across many games or agents. That unification and the stochastic extension are the concrete new pieces; the rest follows from applying existing supermartingale arguments to game payoff matrices. The construction is clean when the benchmark mixed strategies or correlation device are known exactly, which matches the abstract's framing. The soft spot is exactly the one the stress-test flags: everything rests on the benchmark being given in advance. The conditional expectation step that delivers the supermartingale property uses the known equilibrium payoffs directly; if the benchmark instead has to be estimated from the same stream of play, that step fails and the anytime validity and detection-time bounds no longer hold. The paper does not appear to supply a workaround or robustness result for estimated benchmarks, so the practical scope is narrower than the title might suggest. This is for people working on online monitoring in multi-agent systems or sequential testing in games. A reader who already knows e-values and wants to see them applied to equilibrium conditions will get the most out of it. The work is coherent on its own terms and the claims are specific enough to check, so it deserves a serious referee to verify the supermartingale derivations and the stochastic-game extension.","headline":"Applies e-value supermartingales to build anytime-valid tests for deviation from a known benchmark equilibrium, unifying Nash/CCE/CE and extending to stochastic games with target policies.","tokens_in":2360,"tokens_out":437,"would_cite":false,"duration_ms":17730,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Sequential e-value supermartingales for equilibrium deviation detection; no overlap with RS cost or ratio forcing","alignment":"orthogonal","rationale":"Paper constructs supermartingales via per-round factors E=1-λX (X=payoff deviation) and their products/mixtures to accumulate evidence against Nash/CCE benchmarks, with Ville bounds and detection-time analysis. This is standard sequential hypothesis testing machinery. RS framework (reality_from_one_distinction, Jcost uniqueness via washburn_uniqueness_aczel in Cost/FunctionalEquation, phi-ladder constants, 8-tick/D=3 forcing in DimensionForcing/AlexanderDuality) derives specific reciprocal cost J(x)=½(x+x⁻¹)-1 and geometric structures from bare distinguishability; paper neither uses nor contradicts any of these. Domain (cs.GT sequential testing) lies outside RS scope.","tokens_in":61170,"confidence":"high","tokens_out":198,"duration_ms":6733,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A sequential e-value framework detects deviations from strategic equilibria in multi-agent interactions without requiring a fixed sample size.","keywords":["e-values","anytime inference","multi-agent systems","equilibrium detection","supermartingales","repeated games","strategic deviations","stochastic games"],"falsifier":"A simulation or real-world dataset where agents systematically violate equilibrium payoffs but the supermartingale fails to grow significantly, or where matching play triggers false detection.","tokens_in":2662,"feed_emoji":"","tokens_out":467,"duration_ms":18246,"temperature":0.7,"pith_summary":"The paper introduces a method to monitor whether agents in repeated games are sticking to equilibrium play or deviating over time. It uses the e-value approach to build a supermartingale that grows when observed payoffs violate the benchmark conditions. This allows continuous testing that remains valid at any stopping time. Such a tool matters because multi-agent systems like markets or autonomous vehicle fleets can drift from expected rational behavior, and early detection helps maintain stability. The approach covers Nash, correlated, and coarse correlated equilibria in normal-form games and extends to stochastic settings.","feed_headline":"E-value supermartingales detect game equilibrium deviations online","feed_subtitle":"The test accumulates evidence of payoff violations in real time without a preset sample size and works for Nash and correlated equilibria.","key_machinery":"The test supermartingale derived from the e-value framework by betting against the benchmark payoff conditions.","core_discovery":"By betting against a known benchmark equilibrium, the method constructs a test supermartingale that accumulates evidence of departure whenever payoffs systematically violate the equilibrium conditions, providing a statistically valid measure of deviation that can be monitored online with finite-time guarantees.","pith_inferences":["Could apply to monitoring AI agents in multi-agent environments for misalignment.","May integrate with online learning algorithms to trigger interventions upon detected deviations.","Potential use in regulatory oversight of algorithmic trading or automated markets."],"forward_implications":["Unified detection for Nash, correlated, and coarse correlated equilibria in repeated normal-form games.","Finite-time guarantees and detailed analysis of detection times.","Increased detection power in large games via Benjamini-Hochberg procedures while controlling false discovery rate.","Extension to stochastic games to verify adherence to target policies online."],"fun_headline_variants":["E-value supermartingales detect multi-agent equilibrium breaks","Supermartingales monitor deviations from Nash equilibria online","Anytime e-value betting against game equilibrium benchmarks","Framework tracks strategic deviations in repeated normal-form games","Test supermartingales accumulate evidence of payoff violations online"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The benchmark equilibrium is known in advance and observed play can be directly compared to its payoff conditions without extra assumptions on deviation behavior.","fun_headline_variants_meta":{"raw":{"variants":["E-value supermartingales detect multi-agent equilibrium breaks","Supermartingales monitor deviations from Nash equilibria online","Anytime e-value betting against game equilibrium benchmarks","Framework tracks strategic deviations in repeated normal-form games","Test supermartingales accumulate evidence of payoff violations online"]},"model":"grok-4.3","cost_usd":0.010111,"raw_usage":{"total_tokens":4466,"prompt_tokens":628,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":101112000,"prompt_tokens_details":{"text_tokens":628,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3771,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":628,"tokens_out":67,"duration_ms":22110,"temperature":1.0,"reasoning_tokens":3771,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T07:37:39.245592+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A simulation or real-world dataset where agents systematically violate equilibrium payoffs but the supermartingale fails to grow significantly, or where matching play triggers false detection.","supporting_citations":[],"review_version":1}