{"id":"bd50695c-f0fd-4971-aec6-efd529f44d2e","arxiv_id":"2608.08239","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"When a different LLM is substituted into an in-progress agent trajectory, the trajectory diverges almost immediately from the logged one, so replay-based routing benchmarks evaluate decisions against states that never occur.","lead":"This paper tests whether static log-replay benchmarks can evaluate per-step model switching in LLM agents, by forking live SWE-bench trajectories and continuing each fork with a different model in a rebuilt environment. It finds that replayed states become invalid almost immediately after a model swap, and that replay evaluators mispredict every success-relevant outcome in the study.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Outcome-level claim rests on five decisive events and a bespoke log-stitching rule; the abstract's 'mispredicts every' overreaches, though the action-level replay gap is solid.","rationale":"The action-level core of the paper is careful and reproducible: paired bootstrap CIs, control forks, prefix-replay fidelity, six run pairs, and released artifacts. I see no reason to doubt that early per-step swaps diverge far above the noise floor and that replay validity is low; that alone shows the inputs to replay evaluation are invalid in this regime. The stress-test question is whether the abstract's outcome-level claim—'mispredicts every success-relevant outcome call'—is load-bearing and adequately supported. I find it is load-bearing because the title and abstract sell a wrong-world result, not merely invalid inputs, and it is the outcome misprediction that makes the paper's contribution prescriptive for router benchmarks. But the evidence is thin: five decisive calls, a bespoke stitch rule that no published benchmark uses, and a 0–3% success regime in which success-relevant events are rare. The exact binomial CI for 0/5 does not exclude chance-level performance. This does not overturn the action-level result, but it means the strongest version of the central claim should be conditional on more data or on softened language. The paper's own Limitations section already concedes the key points; the mismatch is with the abstract and title. Since the reader's CONDITIONAL verdict already captures this, my recommendation is UNCHANGED. I partially agree with the reader: they emphasized representativeness of the pilot pool; I emphasize the specific fragility of the outcome-level event count and the non-standard evaluator.","tokens_in":8926,"tokens_out":10072,"duration_ms":95820,"concrete_test":"Compute the exact binomial 95% confidence interval for the true success-relevant error rate from the observed 0/5 decisive calls. If the upper bound exceeds 0.5, the data cannot rule out chance-level performance, and the abstract's 'mispredicts every success-relevant outcome call' should be downgraded to a sign-only result pending a larger sample of success-relevant events.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two tiers. The action-level tier (§4.1, Table 1) is well supported: swap arms exceed same-model controls in all 24 cells, bootstrap CIs exclude zero, the AWQ-served reverse direction gives a near-deterministic control floor, and prefix-replay fidelity (99.99%) excludes environment reconstruction error. But the abstract's outcome-level tier—'replay mispredicts every success-relevant outcome call'—is load-bearing for the headline that replay 'scores the wrong world' in the sense of producing wrong benchmark scores, and it is not adequately secured. First, the log-stitching evaluator in §4.5 is defined by the authors, not taken from any published benchmark; the paper explicitly disclaims that any specific benchmark implements precisely this stitching rule. Real per-step routing benchmarks may collect outputs at the decision point from the base prefix rather than reuse the target model's full standalone trajectory, so a 0-for-5 result on this bespoke rule does not directly indict actual evaluation procedures. Second, with base success rates of 0–3%, success-relevant calls are rare; five decisive calls cannot distinguish a systematically broken evaluator from chance—a true success-detection accuracy of 45% would produce 0/5 with probability about 5%, and the exact binomial 95% CI upper bound is near 45%. The Limitations section honestly concedes the small n and explicitly stops short of demonstrating end-to-end router mis-ranking, yet the abstract retains the unqualified phrase. Thus the principal empirical evidence for the strongest version of the claim is statistically fragile and operationally non-standard.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether replay-based static evaluation of per-step model switching in LLM agents is valid. On SWE-bench Verified with a mini-SWE-agent scaffold, the authors fork live trajectories at 30% and 70% of the base trajectory, re-execute the prefix in a fresh container, and continue with either a different model (swap arm) or the same model (control arm). The two models are Qwen3-4B-Instruct (FP8) and Qwen3-14B (AWQ). The action-level comparisons show that swaps diverge from the logged base trajectory far more than same-model controls, with bootstrap CIs excluding zero in all four arms and the swap-above-control ordering holding in all 24 run-pair cells. The paper also reports five outcome flips in swap arms, zero in 359 control forks, and that a log-stitching replay evaluator mispredicts five success-relevant calls and produces near-orthogonal patches. The conclusion is that replay-based benchmarks for agentic routing evaluate decisions against states that never occur.","tokens_in":9161,"tokens_out":6155,"duration_ms":61384,"significance":"The action-level result is a meaningful, well-controlled negative result. Its strengths are substantive: paired bootstrap inference at the instance level, same-model control arms for the noise floor, prefix-replay fidelity checked at 99.99% return-code agreement over 11,702 actions, replication across six run pairs and all 24 cells, and release of the harness and full trajectory dataset. If the structural claim 'replay scores states that never occur' holds, it directly challenges the validity of a growing class of agentic router benchmarks. The evidence is less secure for the outcome-level claims: the success-relevant calls number five, the stitching rule is bespoke to the paper, and the pilot's absolute success rates are 0-3%, so the abstract's wording 'replay mispredicts every success-relevant outcome call' overreaches what the data can show. The paper is a rigorous pilot whose main claims need to be re-scoped rather than discarded.","major_comments":[{"comment":"The claim that replay 'mispredicts every success-relevant outcome call' is not supported by five decisive events. With five Bernoulli trials, an evaluator with true success-detection accuracy as high as 45% would still be wrong on all five calls with probability about 5%, and the exact 95% confidence interval for 0/5 successes has upper bound near 45%. The five events can indicate a sign, but not a systematic 0-for-5 failure. In addition, the stitching rule is defined by the authors and the paper explicitly disclaims that any published benchmark implements it. A real per-step router evaluator might collect the target model's decision at the fork point from the base prefix rather than from the target model's full standalone run. Please downgrade this result to a pilot-level observation, remove 'mispredicts every' from the abstract, or add a binomial analysis that quantifies what the five calls can actually exclude.","section":"§4.5, Table 2, Abstract"},{"comment":"The Limitations section states that the paper shows the inputs to replay evaluation are invalid 'rather than demonstrating end-to-end router mis-ranking,' yet the title and abstract assert generally that replay-based benchmarks 'score the wrong world for agentic routing.' The action-level evidence does establish that many replayed states never occur, but the broader claim about benchmark scores and router rankings goes beyond what the pilot can demonstrate. Please align the title and abstract with the Section 6 scope, or add an explicit bridge explaining why the structural invalidity of replayed states is sufficient for the broader conclusion despite the pilot's narrow regime.","section":"§6, Abstract, Title"},{"comment":"The outcome-flip evidence is reported as 'all five flips in swap arms, zero in 359 control forks,' but these events are selected post hoc from 180 instances across three difficulty tiers and two fork positions, with no multiplicity control. With base success rates of 0-3%, the counts cannot distinguish a genuine swap-only flip tendency from a rare-event null. The fork-position invariance of two instances is a useful case-study observation, but the flip counts should be presented as illustrative rather than as statistical evidence, and they should not appear in the abstract's headline as a stand-alone result.","section":"§4.2"}],"minor_comments":[{"comment":"The abstract says '74-77% of early swaps diverge at the first post-fork action,' but Section 4.1 reports 73.9% (up) and 76.7% (down); please make the numbers consistent.","section":"Abstract vs. §4.1"},{"comment":"The text alternates between 'patch-identity metrics' and 'patch similarity.' Since the reported metric is a SequenceMatcher ratio over raw unified diffs, use 'patch similarity' consistently to avoid implying exact identity is being measured.","section":"§4.4"},{"comment":"The parenthetical explaining that the five decisive calls differ from the five flip events of Section 4.2 is easy to misread. A short worked example of one missed rescue, one false success, and one correctly-called downgrade loss would clarify the overlap and the direction of each error.","section":"§4.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-executed pilot with a strong action-level result and open artifacts. The main barrier is the overclaimed outcome-level conclusion and the general-sounding title/abstract relative to the explicitly narrow regime. If the authors re-scope those claims and add a small binomial or exact-test framing for the five decisive calls, I would be willing to accept the revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe core of this paper is a measurement the agentic-routing field needed: how far a replayed trajectory drifts when you swap models mid-agent, and the branching-rollout protocol with same-model controls is the right way to get it. The action-level result is solid. The authors check prefix-replay fidelity (99.99% return-code agreement), resample at instance level, and replicate across six run pairs; in all 24 cells the swap arm exceeds its control. That early swaps leave only 3% of replayed states valid is a stark, useful number. This alone justifies publication.\n\nThe outcome-level claims are where the paper overreaches. The abstract says replay \"mispredicts every success-relevant outcome call,\" but that rests on five decisive events. With 0-3% base success, five calls cannot separate a broken evaluator from a mediocre one: a detector with ~45% true accuracy would produce 0/5 with probability around 5%, and the binomial CI upper bound is near 45%. The paper's body honestly says it claims only the sign, but the abstract doesn't carry that caution. Also, the log-stitching evaluator is bespoke; the authors themselves disclaim that any published benchmark uses that exact rule. So the direct indictment of existing evaluations is weaker than the headline.\n\nThe limitation section is exemplary: it calls the work a controlled pilot and explicitly stops short of claiming end-to-end router mis-ranking. The title and abstract push past that. The fix is straightforward: qualify the outcome-level sentence and present the action-level replay validity as the robust, generalizable result.\n\nGeneralization is a question, not a flaw. One scaffold, two quantized models from one family, n=30 per pair, a 24GB serving budget. Divergence likely shrinks with competent agents and unquantized pools, but the off-distribution problem does not disappear. The implications section is properly hedged.\n\nThis deserves serious peer review. The protocol is a real methodological contribution, code and data are released, and the action-level measurement is convincing. I'd send it out with a request to soften the abstract and add explicit uncertainty about the outcome statistics. Worth citing, and I'd bring it to a reading group.","headline":"Branching-rollout protocol gives a solid action-level replay-gap measurement; outcome-level claims and abstract overreach, but the paper deserves a serious referee.","tokens_in":9741,"tokens_out":3412,"would_cite":true,"duration_ms":29638,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replay evaluation scores LLM agent model switching against states that never occur.","keywords":["LLM routing","agent evaluation","replay gap","branching rollouts","off-policy evaluation","model switching","counterfactual evaluation","SWE-bench"],"falsifier":"Run the same branching protocol with a competent, unquantized model pair and a base success rate above 20%; if a log-stitching replay evaluator correctly predicts most success-relevant switches, or if swap arms no longer exceed same-model control floors in edit distance, the claim that replay scores the wrong world would be contradicted.","tokens_in":8679,"feed_emoji":"🔀","tokens_out":6477,"duration_ms":55205,"temperature":0.7,"pith_summary":"This paper sets out to show that static replay evaluation of per-step model switching in LLM agents evaluates decisions against states that never occur. On a software-engineering agent benchmark, the authors fork live trajectories at controlled points and continue with a different model, comparing against same-model control forks that isolate noise. Swap arms exceed their control floors by +0.25 to +0.66 in normalized action edit distance, and 74–77% of early swaps diverge at the first post-fork action, leaving only 3% of replayed states valid. A log-stitching replay evaluator mispredicts every success-relevant outcome call, predicting patches with 0.00–0.11 similarity to reality. If this holds, replay-based benchmarks for agentic routing are structurally invalid, and live branching evaluation or corrected off-policy estimators become necessary.","feed_headline":"Static replay judges agent routers on states that never occur","feed_subtitle":"Branching rollouts show swaps rewrite 61-94% of post-fork actions; benchmarks need live evaluation.","key_machinery":"The load-bearing mechanism is the branching-rollout protocol: at a chosen fork step, the recorded action prefix is re-executed in a fresh container, the message history is seeded with the recorded prefix, and the trajectory continues with a different (or, for control, the same) model. The comparison between swap arms and same-model control arms isolates sampling and serving-stack noise, so divergence attributable to the swap is read relative to the control. Three metrics carry the result: normalized action edit distance between post-fork suffixes, replay validity (the prefix-match fraction of base post-fork actions), and patch similarity between the replay-predicted and actually-produced patches.","core_discovery":"The central claim is that the assumption underlying replay evaluation of agentic routers—that substituting another model's recorded output at step k leaves the rest of the trajectory unaffected—is wrong in a measurable, structured way. Using branching rollouts on SWE-bench with a bash-only agent scaffold, the paper finds that model swaps rewrite 61–94% of post-fork actions, with first-action divergence in up to 77% of early swaps. Replay validity, the fraction of post-fork states a replay evaluator would score correctly, falls to 3–8% for early swaps; a replay evaluator of an early swap scores 92–97% of post-fork decisions against a state that never occurs. All five observed outcome flips occur in swap arms, with none in 359 control forks, and a log-stitching replay evaluator mispredicts every success-relevant outcome while achieving only 0.00–0.11 patch similarity.","pith_inferences":["The structural invalidity likely extends beyond bash coding agents to any closed-loop environment—web navigation, computer use, tool use—where the next observation depends on the previous action; the paper tests one scaffold, so this is an extrapolation.","A testable extension: with a third fork position at 50% and a competent unquantized model pair, the depth-dependence of divergence should saturate or invert, which would pin down where handoffs are safe.","The paper's 'thoroughness tax' observation suggests that upgrading under a fixed step budget can reduce completion rate; if it generalizes, routing cost models must include expected step consumption, otherwise reported efficiency gains will be overstated."],"forward_implications":["Replay-based benchmarks for agentic routing that look up logged outputs cannot observe outcome flips; any evaluation of per-step switching needs live branching rollouts.","Early swaps at 30% of a trajectory are the most invalid for replay evaluation, while late handoffs at 70% inherit more completed work and diverge less.","Upgrades diverge immediately because the stronger model re-decides at once, whereas downgrades diverge later and less, so handoff policy should depend on direction and position.","Under tight budgets, the stronger model may exhaust its step limit without submitting more often than the weak one, so routers should price step consumption, not only per-step quality.","Off-policy estimators such as importance sampling and doubly robust estimation need branched ground truth to validate them, because per-step action-probability ratios are far from one immediately after a swap."],"supporting_citations":[{"why":"Provides the SWE-bench Verified task set and per-instance Docker environments used for forking.","marker":"Jimenez et al., 2024"},{"why":"Supplies the minimal bash-only ReAct scaffold whose trajectories are forked in the branching protocol.","marker":"Lieret et al., 2025"},{"why":"Exemplifies the static replay evaluation design this paper argues is invalid for agents.","marker":"Hu et al., 2024"},{"why":"Supplies the importance-sampling framework that motivates the need for branched ground truth in off-policy evaluation.","marker":"Precup et al., 2000"},{"why":"Defines the model family from which the two swapped deployment configurations are drawn.","marker":"Yang et al., 2025"}],"fun_headline_variants":["Replay benchmarks score agent routers on ghost states","Static replay misses 94% of post-fork actions","Agent routing replay rewrites 61–94% of actions","Replay eval sees a world that never occurs","Log replay flips every outcome it scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pilot pool—one bash-only scaffold, two quantized models from one family, 30 instances per run pair, and 0–3% base success rates—is representative of agentic routing generally, so that the measured divergence and outcome misprediction transfer to competent, larger-scale regimes.","fun_headline_variants_meta":{"raw":{"variants":["Replay benchmarks score agent routers on ghost states","Static replay misses 94% of post-fork actions","Agent routing replay rewrites 61–94% of actions","Replay eval sees a world that never occurs","Log replay flips every outcome it scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000532,"raw_usage":{"total_tokens":2626,"prompt_tokens":1074,"completion_tokens":1552,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":690,"completion_tokens_details":{"reasoning_tokens":1477}},"tokens_in":690,"tokens_out":1552,"duration_ms":12618,"temperature":1.0,"reasoning_tokens":1477,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:14:08.584863+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same branching protocol with a competent, unquantized model pair and a base success rate above 20%; if a log-stitching replay evaluator correctly predicts most success-relevant switches, or if swap arms no longer exceed same-model control floors in edit distance, the claim that replay scores the wrong world would be contradicted.","supporting_citations":[],"review_version":1}