{"id":"68485345-ff8d-4bbb-8630-51f56ac8fa65","arxiv_id":"2608.07911","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Trace-driven MoE expert-cache results are fragile to replay semantics, prompt-template contamination, and operating regime; after correction, the apparent offline headroom is mostly not recoverable by lightweight causal predictors.","lead":"This paper shows that common choices in replaying and preparing traffic traces for Mixture-of-Experts cache experiments can flip which cache policy looks best. It proposes concrete measurement fixes and shows that, after correction, the gap to an ideal cache mostly depends on information that simple learned predictors cannot recover.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative claims rest on one 4-bit 128-expert model and one replay contract; the frontier-scale regime that motivates the paper is the one the data covers least well, leaving the portability of the central claim unestablished.","rationale":"The reader's conditional verdict is well aligned with the actual risk profile. The paper is internally careful, extensively self-audited, and explicitly narrow in its claims; the headline assertion that a large offline-optimal gap overstates the gains recovered by representative lightweight causal mechanisms is supported by the data as presented. The soft spot is not a flaw in the argument's logic but in the reach of its evidence. The strongest quantitative support comes from one 4-bit, 128-expert model, and the paper itself states that the regime it most wants to inform is the one its data cover least well (§11). Because the entire measurement chain begins with router decisions from that quantized model, full-precision routing differences could propagate into the gap, the decomposition, and the predictor's negative result. The replay contract is a second scoping condition: the event-atomic model is a legitimate contract, but so is serialized loading (§2.2), and all headline numbers are computed under the former. I therefore do not see grounds to reject the paper or to demand a different verdict than the reader's conditional one, but the condition should explicitly include verification that the central numbers survive full-precision or larger-model routing traces and at least one alternative replay contract. This is why the verdict should remain unchanged rather than be moved to unconditional acceptance.","tokens_in":30171,"tokens_out":13977,"duration_ms":187711,"concrete_test":"Re-run the §7 pipeline on the same Qwen3-30B-A3B at full precision, or on an open MoE with at least 256 experts and top-16 routing, using identical prompts, batching, ρ=40%, and B=8, and recompute Tables 12–14 (gap, forced-admission decomposition, predictor recovery). Independently, implement a serialized per-access replay on the same event stream and capacity to measure how much of the 44–46% gap and the 84–97% future-victim share is contract-specific. If the gap or the decomposition moves by more than about 10 relative points, or the predictor's sign flips, the central claim must be re-scoped to the measured model and replay contract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (§7.4) is deliberately scoped to 'our evaluated settings,' but all primary numbers—the 44.2–45.9% stable gap, the 84.3–96.6% future-victim share, and the −11.4% predictor recovery—are computed on one 4-bit quantized Qwen3-30B-A3B with greedy decoding. §11 concedes both that quantization may change routing and that the frontier-scale regime motivating the work is the one the collected traces cover least well. The results are also computed under a single fused-event replay contract (Definition 2), while §2.2 explicitly accepts serialized intra-layer loading as a legitimate alternative. If router behavior differs under full precision, at larger expert counts, or under a serialized execution contract, the measured gap, its decomposition, and the predictor failure could all shift. This is not an internal inconsistency; it is an external-validity gap, but it is load-bearing because the paper's practical message is that reporting an offline gap without decomposition invites an unwarranted inference for MoE cache evaluations generally, not only for this model and contract.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies trace-driven evaluation of MoE expert caching and argues that three evaluation choices—replay semantics, workload probe construction, and operating-regime alignment—can change conclusions rather than merely shift numbers. Under an event-atomic replay contract it shows that flattened per-access replay inflates recency-based policies by 27–29% and inverts policy rankings. It demonstrates that single-template workload probes generate verbatim shared prefixes that inflate apparent expert locality, and it introduces a matched-pair rendering control plus diversity-controlled probe set. It shows that normalized miss fractions do not transfer across models and that the per-step expert-union-to-capacity ratio is necessary but not sufficient, since temporal reordering of an identical event stream changes the offline-optimal gap. After correcting these axes, a stable offline-optimal gap of 44.2–45.9% remains at rho=40%, B=8; a forced-admission oracle attributes 84.3–96.6% of this gap to future-victim knowledge, and a linear next-use predictor recovers -11.4% of the gap. The paper closes with a reporting checklist, four scoped negative results, one invalidated design, and a detailed self-audit. The central claim is deliberately narrow: in the authors' evaluated settings, a large offline-optimal gap substantially overstates the gains recovered by representative lightweight causal mechanisms.","tokens_in":30380,"tokens_out":6030,"duration_ms":74348,"significance":"If the results hold, the paper makes a strong methodological contribution. The machine-checkable reference trace, pre-registered confirmatory thresholds, disjoint discovery/confirmatory splits, matched controls, and candid disclosure of invalidated designs are exemplary and raise the bar for this literature. The forced-admission decomposition of the offline gap into bypass-admission and future-victim components is a genuinely useful diagnostic, and the finding that a next-use predictor with positive transfer R2 is anti-correlated with true victim ranking at decision points is a surprising, falsifiable result. The paper's three scoped negative results and the invalidated partition experiment are reported with unusual honesty. The main limitation is external scope: the primary quantitative claims rest on one 4-bit quantized 128-expert model and one replay contract, and the paper itself concedes that the frontier-scale regime motivating the work is the one its data cover least well. This does not undermine the internal logic, but it does bound how far the headline conclusion can be generalized without additional evidence.","major_comments":[{"comment":"The central decomposition and predictor-failure claim rest on a single 4-bit quantized 128-expert model (Qwen3-30B-A3B) with greedy decoding, under one replay contract. Section 11 concedes both that quantization may change routing and that no trace was collected at the 896-expert scale that motivates the application. Since the paper's practical message extends beyond this one model—'reporting such a gap without decomposing it invites an unwarranted inference'—the external-validity gap is load-bearing. I ask for either a second full-precision or otherwise non-quantized trace at comparable expert count with the same decomposition, or an explicit scoping of the §7.4 claim to the 128-expert 4-bit fused-event setting in the abstract and conclusion, with the checklist presented as the only generalizable output.","section":"§7.4, §11"},{"comment":"The decomposition is defined only for Definition 2's fused-event contract, yet §2.2 explicitly accepts serialized intra-layer loading as a legitimate alternative execution model. The 27–29% replay inflation in Table 2 demonstrates sensitivity to semantics, but the paper never quantifies how m_c, m_B, m_F, or the 84.3–96.6% future-victim share would change under a serialized contract. If serialized execution is legitimate, then the headline statement that most of the gap is future-victim knowledge is contract-conditional. Please add a sensitivity analysis for a serialized intra-event replay, or state in the central claim that all §7 numbers are conditional on the fused-event atomic contract.","section":"§2.2, §7.2"},{"comment":"The causal predictor is a single linear model over eight features, and the negative result is presented as representative of 'representative lightweight causal mechanisms.' Table 16's central exhibit—anti-correlated victim ranking at decision points—is striking, but its force depends on whether the predictor class is representative of what a practitioner would try. The paper acknowledges that a predictor with materially higher transfer accuracy could change the result, but it does not include any stronger baseline on the same features. I request either a second predictor class (for example, gradient-boosted trees or a small MLP with the same features) or a consistent rewording of the abstract and conclusion to say 'a linear next-use predictor fails' rather than 'representative lightweight causal mechanisms fail.'","section":"§7.3, Table 16"}],"minor_comments":[{"comment":"The text contains repeated typographical artifacts such as 'traﬀic' and 'suﬀicient' with nonstandard ligatures; these should be normalized to standard 'traffic' and 'sufficient' in the final version.","section":"Throughout"},{"comment":"The section numbering jumps from §10.4 to §10.6, and Appendix B.2 refers to 'the reporting block of §10.5,' which has no heading in the main text. Add the missing §10.5 heading or renumber the checklist subsections consistently.","section":"§10, Appendix B"},{"comment":"Table 8 is badly corrupted in the supplied text: the tool agent row and the interval notation are unreadable as printed, and Table B1 contains a garbled row label. These tables must be repaired before publication because they support the workload-contamination and static-vs-dynamic claims.","section":"Table 8, Table B1"},{"comment":"The text says the static identity predicts a miss ratio of 1−0.723 = 27.99% of union accesses, but 1−0.723 = 27.7%. If the simulator's 27.99% comes from unrounded shares, please state that explicitly; otherwise correct the arithmetic.","section":"§9.2"}],"recommendation":"major_revision","confidential_remarks":"This is an unusually careful paper, and the self-audit with pre-registration documents is a model for the field. My main concern is that the narrow model/contract evidence base may be over-read by citing papers if the scoping language in §7.4 is not retained or strengthened. I would encourage the editor to require the authors to keep the 'in our evaluated settings' qualification in the abstract and conclusion, and to treat the checklist as the portable contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this one. The paper's core deliverable is not a new caching policy but a diagnosis of why trace-driven MoE cache evaluations mislead. Three constructs stand out: event-atomic replay for the fused-event contract, a matched-pair prompt-rendering control with contamination diagnostics, and the union-to-capacity ratio r_bar as a regime variable. The forced-admission decomposition of the offline-optimal gap is also clean. The reporting checklist in §10 is the most portable output, and it earns its keep: the paper applies it to itself, discloses its own invalidated designs, and releases a hand-verifiable reference trace. That level of self-auditing is rare and genuinely useful.\n\nThe internal methodology is unusually strong. Pre-registered thresholds, disjoint discovery/confirmatory splits, tie-seed sensitivity, and a clear distinction between implementation-stable and statistically-bounded numbers. The central claim—that a large offline-optimal gap overstates what lightweight causal mechanisms recover—is supported by the tables, and the decomposition attributes 84–97% of the gap to future-victim knowledge. The next-use predictor failing worse than random at victim selection is a striking result, scoped sensibly.\n\nThe soft spots are external-validity gaps, not internal errors. The primary quantitative claims rest on one 4-bit 128-expert model with greedy decoding. The frontier-scale regime that motivates the paper is the one the traces cover least well, and the paper admits this. The replay contract is the fused-event one; the paper explicitly accepts serialized intra-layer loading as legitimate, and the measured inflation and decomposition could shift under that alternative. And the artifacts—simulator, probe set, traces—are only promised on publication, which limits immediate independent verification. None of these undermine the internal logic; they cap how far the results generalise. The paper itself is admirably clear about that.\n\nWho benefits? Anyone evaluating MoE caching or, more broadly, doing trace-driven cache studies. The checklist and diagnostics are worth adopting even if the specific numbers don't transfer. I'd send it to a serious referee, with a request that the author release the artifacts and ideally add a second model or a hardware check. The claim is narrow enough that the lack of frontier-scale data is a limitation, not a fatal flaw.","headline":"A genuinely careful measurement-fragility study with a portable checklist; the headline numbers rest on one 4-bit model, but the paper scopes its claims honestly and deserves a real referee.","tokens_in":30913,"tokens_out":1280,"would_cite":true,"duration_ms":16433,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Trace-driven evaluation of MoE expert caching is fragile in three specific ways; after correction, the remaining 44–46% offline-optimal gap is mostly future-victim knowledge that a representative causal predictor cannot supply.","keywords":["mixture-of-experts","expert caching","trace-driven evaluation","replay semantics","workload contamination","offline-optimal gap","next-use prediction","cache eviction"],"falsifier":"Collect a decode-routing trace from a frontier-scale model (896 routed experts, top-16) at a batch size where $r_{\\mathrm{bar}} \\approx 0.3$, replay it under the paper's event-atomic protocol, and compute the forced-admission decomposition. If the future-victim share is well below 84% there, or if a next-use predictor trained on that trace recovers a positive fraction of the gap, the paper's central claim would fail. Separately, an end-to-end engine that serializes expert loads within a layer and shows no 27–29% inflation for LRU under per-access replay would falsify the replay-semantics axis for that execution contract.","tokens_in":29930,"feed_emoji":"🧠","tokens_out":10259,"duration_ms":99497,"temperature":0.7,"pith_summary":"The paper tries to establish that the standard way of measuring how much an expert-cache policy can help MoE inference is fragile enough to produce reversed conclusions, and that the honest remainder is mostly not actionable. Three evaluation choices are isolated: replay semantics (whether a fused (step, layer) cache event is replayed as individual accesses), workload construction (whether probe categories share instruction templates), and operating regime (the ratio of per-step expert union to per-layer capacity). Each axis is shown to change conclusions, not just numbers: sequential replay inflates recency-based policies by 27–29% and inverts their ranking; single-template probe sets manufacture apparent workload locality; aligning models on batch size instead of the union-to-capacity ratio hides a 36.8-percentage-point spread. After all corrections, a stable 44.2–45.9% gap to the offline optimum remains at rho = 40% and B = 8, but the forced-admission decomposition attributes 84.3–96.6% of it to knowing which resident expert is used furthest in the future. A linear next-use predictor fitted on causal features recovers -11.4% of the gap, choosing an optimal victim barely more often than random, which supports the paper's narrow central claim: reporting a large offline-optimal gap without decomposing it overstates the gains recoverable by lightweight causal mechanisms.","feed_headline":"Why MoE cache headroom is mostly unreachable","feed_subtitle":"Three evaluation fixes reveal 84–97% of the MoE cache gap needs impossible future knowledge; a causal predictor fails.","key_machinery":"The argument is carried by four instruments. (1) Event-atomic replay (Definition 2): the cache treats the batch-wide expert union $E(s,l)$ at each (step, layer) as one committed unit, classifies all members against a start-of-event snapshot, and applies admission and eviction only at the event boundary; this is what makes sequential replay's intra-event eviction visible as an artifact. (2) A matched-pair probe design in which the same source records are rendered with diverse versus fixed templates, turning prompt-surface contamination into a measurable treatment. (3) The regime ratio $r(s,l) = |E(s,l)|/c_l$, the per-step per-layer expert union divided by the per-layer cache quota, with the claim that its distribution must be reported for cross-model comparison even though its mean is not sufficient. (4) The Belady-forced-admit decomposition, which splits the offline-optimal gap into a bypass-admission share and a future-victim-selection share, plus a causal next-use-distance predictor substituted into the same eviction and admission machinery.","core_discovery":"Under an event-atomic replay protocol—where the expert set touched by the whole batch at each (step, layer) is one committed unit, hits and misses are classified against the snapshot at event start, and admission/eviction is deferred to the event boundary—the paper finds that the offline-optimal gap to Belady is large and stable across workload compositions, but that most of it is not reachable by the causal mechanisms a practitioner would use. The decomposition via Belady-forced-admit shows bypass admission accounts for 15.7% at B = 8 and 3.4% at B = 2 of the gap, leaving 84.3% and 96.6% to future-victim selection. A next-use-distance predictor with strictly more features than the best causal baseline transfers at $R^2 = 0.24$ and still performs worse than LFRU, recovering -11.4% of the gap; at eviction decision points it picks an optimal victim 3.39% of the time versus 2.42% for random and 20.6–22.1% for LRU and LFRU. The paper states the claim narrowly: a large offline-optimal gap substantially overstates the gains recovered by representative lightweight causal mechanisms in these evaluated settings, so a gap reported without the forced-admission decomposition invites an unwarranted inference.","pith_inferences":["The checklist's likely effect, if adopted, is to make a raw offline-optimal gap alone an insufficient motivation for a cache controller; future work would need to name the recoverable share, which should push effort toward admission policies, prefetching, and scheduling rather than generic eviction learning.","The union-to-capacity ratio suggests a testable scaling rule: the recoverable gap should collapse as $r_{\\mathrm{bar}}$ crosses 1, so workloads with larger batches or smaller per-layer caches will have less real headroom; a frontier-scale trace would let someone check whether this holds at 896 experts.","The paper's negative next-use result is for victim ranking, not for the different task of predicting which experts will be routed next; a prefetch policy aimed at the latter could in principle capture part of the gap that eviction cannot, and is a natural next experiment.","If real engines with serialized expert loads within a layer do not show the replay inflation, then the replay-semantics axis reduces to a simulator-correctness requirement rather than a system property; measuring transferred block counts on such an engine would settle the scope."],"forward_implications":["Any simulator that claims fused-event accounting but replays per-access eviction will inflate LRU and LFRU by 27–29% and invert policy rankings; such results should be re-checked under event-atomic replay before being used to motivate a policy.","Cross-model MoE cache comparisons should align on the per-step union-to-capacity ratio $r_{\\mathrm{bar}}$ rather than batch size; the three models here agree to within 1.1–4.6 percentage points when aligned on $r_{\\mathrm{bar}}$, versus a 36.8-percentage-point spread on batch size.","Report the forced-admission decomposition alongside any offline-optimal gap: at the two held-out operating points 84.3–96.6% of the gap is future-victim knowledge, so the raw gap is not available headroom.","Workload-conditioned routing studies should use multiple template forms and matched-pair controls; single-template probe sets manufactured an apparent 57.2% cache benefit that fell to 5.9% after correction and reversed the category ordering.","The reported gap is regime-specific: the 44.18–45.93% range applies only to rho = 40%, B = 8, and must not be pooled with other operating points."],"supporting_citations":[{"why":"Defines the offline-optimal replacement algorithm with bypass that the paper uses to compute the gap and the forced-admission decomposition.","marker":"Belady 1966"},{"why":"Provides the working-set model that frames the resident-set, capacity, and locality arguments.","marker":"Denning 1968"},{"why":"Introduces the sparsely gated MoE layer whose routing behavior the recorded traces capture.","marker":"Shazeer et al. 2017"},{"why":"An end-to-end offloading system that grounds the motivating scenario and delimits where replay artifacts do not apply.","marker":"Eliseev and Mazur 2023"},{"why":"Representative expert-prefetching work whose prediction target (next routed experts) the paper contrasts with next-use-distance prediction.","marker":"Du et al. 2024"},{"why":"Representative expert-prefetching work used to separate the paper's negative next-use result from the adjacent prefetching literature.","marker":"Fang et al. 2025"},{"why":"Supplies C4 input sequences used in the false-positive contamination check of the route-side diagnostics.","marker":"Raffel et al. 2020"}],"fun_headline_variants":["Most MoE cache headroom is unreachable in practice","MoE cache gap is 84–97% future-victim selection","Trace replay flips MoE cache policy winners","Offline-optimal MoE cache gap is a causal mirage","Causal predictors recover negative MoE cache gap"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that MoE serving obeys the fused-event traffic contract—each (step, layer) sees the batch-wide expert union as one committed unit with retention deferred to the event boundary—and that traces from a 4-bit 128-expert model plus two smaller models represent the operating regime where expert caching will be deployed; if a real engine serializes expert loads inside a layer, the replay-semantics results do not apply, and the paper concedes the frontier-scale regime is the one its data covers least well.","fun_headline_variants_meta":{"raw":{"variants":["Most MoE cache headroom is unreachable in practice","MoE cache gap is 84–97% future-victim selection","Trace replay flips MoE cache policy winners","Offline-optimal MoE cache gap is a causal mirage","Causal predictors recover negative MoE cache gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000688,"raw_usage":{"total_tokens":3256,"prompt_tokens":1224,"completion_tokens":2032,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":840,"completion_tokens_details":{"reasoning_tokens":1949}},"tokens_in":840,"tokens_out":2032,"duration_ms":17102,"temperature":1.0,"reasoning_tokens":1949,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:41:29.423626+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a decode-routing trace from a frontier-scale model (896 routed experts, top-16) at a batch size where $r_{\\mathrm{bar}} \\approx 0.3$, replay it under the paper's event-atomic protocol, and compute the forced-admission decomposition. If the future-victim share is well below 84% there, or if a next-use predictor trained on that trace recovers a positive fraction of the gap, the paper's central claim would fail. Separately, an end-to-end engine that serializes expert loads within a layer and shows no 27–29% inflation for LRU under per-access replay would falsify the replay-semantics axis for that execution contract.","supporting_citations":[],"review_version":1}