{"id":"9db03f32-0969-438c-89ed-92947c859481","arxiv_id":"2608.08650","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey proposes eight architectural milestones and four control planes for MoE LLMs, arguing the field is moving toward decoupling routing, compute budgets, and physical execution.","lead":"This paper is a technical survey of Mixture-of-Experts (MoE) architectures for large language models, organizing the history into eight milestones and four control planes. It argues that the main trend is the decoupling of semantic routing, computational budgets, and physical execution, which is a useful map for researchers and engineers navigating MoE design choices.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Nodes 7–8 and the claimed decoupling trend rest on citations with placeholder authors and unverifiable arXiv IDs; if those sources do not exist or misreport the architectures, the central claim loses its empirical support.","rationale":"The paper proposes a coherent organizing framework: four control planes (Topology, Routing, Balance, Expert Parallel) and a dependency graph of eight milestones, with a clear distinction between the historical mainline (nodes 1–6) and the two orthogonal branches (nodes 7–8). The formalization in Eqs. (1)–(7) is standard, and the paper is careful to hedge predictive claims (e.g., Section 3.8 treats node 8 as a research frontier). If the cited systems are real, the framework is a useful synthesis. The load-bearing weakness is not the framework's internal logic but the empirical identity of its distinctive branches. Node 7 (dynamic compute) and node 8 (semantic/physical decoupling) are exactly what distinguishes this survey's 'main trend' claim from a generic history of MoE, and exactly the nodes whose citations are unverifiable placeholders or future-dated arXiv IDs. The paper's own acknowledgment that it was 'completed with the assistance of Codex' and the presence of a duplicated Chinese appendix (unstyled, with its own reference numbering) further support the need for source verification. The reader's verdict of CONDITIONAL is therefore correct, but the primary condition should be explicit: verify the existence and content of the load-bearing references for nodes 7–8, not merely 'address citation issues.' The abstract's promise of 'equal-budget pretraining experiments' should also be corrected to 'experimental design' since Section 9 only proposes a matrix. If the references fail verification, the appropriate outcome would be a rejection or an explicit downgrade of nodes 7–8 to 'hypothetical directions.' I agree only partially with the reader's weakest assumption: third-party accuracy is a fair concern, but the sharper and more decisive issue is the verifiability of the primary sources themselves.","tokens_in":26712,"tokens_out":11527,"duration_ms":109418,"concrete_test":"Resolve references [19], [21], [22], [23], [24], and [26] against arXiv, the ACL Anthology, and the given URL. Concretely: (1) query the arXiv API for 2509.01322, 2602.04870, and 2604.19835 and check that the papers exist and describe the claimed architectures (zero-computation experts for LongCat-Flash; Multi-Head LatentMoE/Head Parallel with communication O(1) in k); (2) open https://aclanthology.org/2026.acl-industry.20/ and https://aclanthology.org/2026.acl-long.2065/ and confirm real author names and matching content; (3) fetch https://longcat.ai/blog/longcat-2.0/ and verify a 1.6T per-core parallel dense/MoE system is described. If any identifier is invalid or the content differs materially, nodes 7–8 and the decoupling trend claim lack documented support and the survey must be revised or reclassified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that the main MoE trend is 'decoupling semantic routing, computational budgets, and physical execution' (Section 11)—is carried by Section 3.7 (node 7: dynamic compute) and Section 3.8 (node 8: semantic/physical decoupling). These sections rest on references [19], [21], [22], [23], [24], and [26]. Several are unverifiable as cited: [22] and [23] list only 'MoHGE Authors' and 'GMoE Authors' as authors, [26] lists 'Expert Upcycling Authors', [19] and [21] attribute models to 'LongCat Team', and [24] is an arXiv ID (2602.04870) with no verifiable authors. The 'LongCat 2.0 Technical Blog' URL is the sole source for a 1.6T-parameter per-core parallel dense/MoE system. The paper itself concedes (Sections 3.8 and 8) that these branches lack public validation at trillion-parameter scale and should be treated as a research frontier, not a mainstream replacement. If these citations are non-existent or the described mechanisms (e.g., zero-computation experts; Head Parallel with communication O(1) in k) are not actually reported there, then nodes 7 and 8 fail the paper's own milestone criterion that later architectures inherit the change (Section 3), and the claimed macro-trend is a speculative projection rather than an observed evolution. This is a sharper failure than the reader's generic third-party-accuracy concern: it questions whether the primary evidence for the two orthogonal branches exists at all.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical survey proposes a framework for organizing the evolution of Mixture-of-Experts (MoE) architectures in large language models. It introduces eight architectural milestones arranged as a dependency graph (six mainline developments, two orthogonal branches) and four control planes (Expert Topology, Routing, Balance, Expert Parallelism). The paper formalizes token-choice MoE, the capacity/quality/system-efficiency trade-off, and the closed-loop relations among the control planes in Equations (1) through (7). It applies this framework to representative models and concludes that the main trend is a shift from activating more sparse parameters toward decoupling semantic routing, computational budgets, and physical execution.","tokens_in":27046,"tokens_out":5833,"duration_ms":58488,"significance":"If the framework is accepted, it offers a useful organizing schema for comparing MoE designs and tracing bottleneck migration. The mathematical formalizations are standard and correct, and the paper credibly separates algorithmic issues (routing, granularity, sharing) from systems issues (all-to-all communication, expert placement, overlap). Concrete strengths include a comprehensive architectural comparison table (Table 6), a practical equal-budget ablation design (Table 7), and a clear articulation of open research questions. However, the manuscript's empirical support is limited: Section 9 contains only a proposed experimental matrix with no results, and the macro-trend claim in Section 11 rests substantially on sources that are either unverifiable or explicitly acknowledged as lacking public validation. These issues affect the paper's central claim rather than peripheral presentation.","major_comments":[{"comment":"The abstract states 'We conclude with equal-budget pretraining experiments,' but Section 9 presents only a proposed experimental design (Table 7) and reports no experimental outcomes. The phrasing in the abstract implies that actual equal-budget experiments were run. Either conduct and report the experiments, or revise the abstract and Section 11 to state that the paper provides an experimental design for future work. This is load-bearing because the macro-trend conclusion is framed as an observed evolution rather than a hypothesis.","section":"Abstract and Section 9"},{"comment":"The central claim that the main trend is decoupling semantic routing, computational budgets, and physical execution rests on nodes 7 and 8. These nodes are supported by references [19], [21], [22], [23], [24], and [26]. Several of these citations are unverifiable as given: [22] lists 'MoHGE Authors' as author, [23] lists 'GMoE Authors', [26] lists 'Expert Upcycling Authors', and [21] is a single blog URL with no archival record. The paper itself concedes in Section 3.8 and Section 8 that these branches lack public validation at trillion-parameter scale and should be treated as a research frontier. Given this concession, the strong conclusion in Section 11 ('the main trend is a shift...') is not supported by the evidence presented. Please either verify and properly archive these sources, or reframe the claim as a projection based on frontier research directions.","section":"Sections 3.7, 3.8, 8, 11 and Table 6"},{"comment":"The paper defines a milestone as a change that satisfies three criteria, the third being that 'later architectures inherit the change.' However, nodes 7 and 8 are explicitly described as orthogonal branches with no clear evidence of inheritance, and the paper itself says they are not 'generations' following node 6. This creates an internal inconsistency: nodes 7 and 8 are called milestones while failing the paper's own inheritance criterion. Please adjust the criteria to accommodate branches (e.g., 'influence future design' rather than 'inheritance'), or rename the components to distinguish established milestones from frontier branches.","section":"Section 3, milestone criteria"}],"minor_comments":[{"comment":"The manuscript contains numerous typographical errors and inconsistent spelling, e.g., 'oﬀicial' in the Abstract, 'eﬀicient' and 'suﬀicient' in multiple sections, and inconsistent use of 'Chapter 3' versus 'Section 3'.","section":"Throughout"},{"comment":"The opening sentence of Section 9 is an incomplete fragment: 'cross-model benchmark cannot identify architectural contributions.' Please provide a subject and connect it to the following sentence.","section":"Section 9"},{"comment":"Several references are non-archival or use placeholder-like author names. For example, [14] and [21] are cited as 'official technical blog' and 'Technical Blog' with only a URL; [22], [23], and [26] use 'Authors' as the author field. Please add access dates, archive links (e.g., DOI or persistent repository), and full author lists where available.","section":"References"},{"comment":"The bias update rule b_i ← b_i + η sign(n̄ − n_i) is described as non-gradient, but it is not clear how η is selected or whether this update interacts with the routing gradient through the Top-k selection. A one-sentence clarification of the design rationale and stability considerations would be helpful.","section":"Equation (7), Section 6.3"},{"comment":"The manuscript includes a full Chinese translation of the paper after the reference list. This duplication is unusual for a journal submission and should be moved to supplementary material or removed, as it distracts from the main text.","section":"Appendix"}],"recommendation":"major_revision","confidential_remarks":"This is a survey rather than a primary experimental paper, and the framework has value as a synthesis. The main risk is the reliance on unverifiable sources for nodes 7 and 8, which carry the macro-trend claim. I recommend that the editor ask the authors to verify that references [19], [21], [22], [23], [24], and [26] exist and are accurately described, or to soften the central claim accordingly. The abstract's promise of actual experiments should also be corrected to avoid misleading readers. The paper's own caveats in Sections 3.8 and 8 make this a fixable issue, so I support major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful survey framework, not a breakthrough, and its main risk is not the math but the sourcing of the two most interesting branches.\n\nWhat's actually new: the eight-milestone dependency graph (six mainline nodes, two orthogonal branches) with a bottleneck-migration criterion for what counts as a milestone, and the four control planes — Topology, Routing, Balance, Expert Parallelism — drawn as a closed loop in Eq. 5. That is a real organizing device, not a reshuffle of prior surveys. I'd also steal two points: the distinction between total All-to-All time and exposed communication time, and the argument that balance objectives should be separated from the LM gradient (the ALF discussion correctly corrects the myth that DeepSeek-V3 has no balance mechanism at all). Section 6 and the proposed ablation matrix in Section 9 are well thought out.\n\nSoft spots, in order. First, the abstract says 'we conclude with equal-budget pretraining experiments,' but Section 9 is only a design proposal; no experiments or data appear anywhere. The reader flagged this and is right; it's an easy rewrite. Second, and more serious: nodes 7 and 8 carry the headline claim that the main trend is decoupling semantic routing from physical execution, and those sections rest on citations I cannot verify — 'MoHGE Authors,' 'GMoE Authors,' 'Expert Upcycling Authors,' a 'LongCat 2.0 Technical Blog' as the sole source for a 1.6T-parameter per-core parallel system, and two 2026 arXiv IDs that a reader cannot check from the text. I can't tell whether these exist. The paper honestly concedes these branches lack public validation and calls them research frontier, but a survey still needs checkable sources for its crown-jewel claims; a referee should run these down. Third, minor: the taxonomy is built from the same systems it explains. That is inherent to survey work, and the explicit milestone criteria mitigate it, so I don't count it as a real flaw.\n\nThe math is standard and correct as far as it goes; Eqs. 1–7 are textbook MoE formalization with no fitted parameters. The mainline citations (Shazeer, GShard, Switch, DeepSeek-V2/V3, Qwen3, Kimi K2) are real and used carefully, including an explicit warning against reading capability numbers as causal comparisons.\n\nWho this is for: researchers designing or comparing MoE pretraining runs, and anyone writing the next MoE survey. It deserves a serious referee. If I were the editor I'd send it out with instructions to verify the node 7/8 references, reword the abstract, and soften the Section 11 trend sentence so it distinguishes the observed mainline from the projected frontier. With those fixes, it is a solid, citable contribution.","headline":"Useful MoE organizing framework, but the headline decoupling trend leans on unverifiable frontier citations and an abstract that promises experiments it doesn't run.","tokens_in":27520,"tokens_out":5155,"would_cite":true,"duration_ms":50806,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that Mixture-of-Experts language-model evolution is best read as a dependency graph of eight milestones and a closed loop of four control planes, whose main trend is decoupling semantic routing, computational budgets, and…","keywords":["Mixture-of-Experts","large language models","sparse routing","load balancing","expert parallelism","dynamic computation","semantic routing","architecture evolution"],"falsifier":"Run a controlled equal-budget pretraining sweep, same data, token budget, and active parameters, varying expert granularity, shared path, balance scope, and execution topology independently, and measure validation loss, specialization, exposed communication, and tail latency; if no control-plane knob changes the outcomes, the bottleneck-migration story would be refuted.","tokens_in":26490,"feed_emoji":"🧩","tokens_out":6958,"duration_ms":69210,"temperature":0.7,"pith_summary":"Mixture-of-Experts (MoE) language models grow parameter counts without growing per-token compute, but the paper argues their history is not a list of bigger models. It organizes the evolution as eight milestones, six mainline and two orthogonal branches, defined by which bottleneck each change removes and where the bottleneck moves next. It then dissects any individual system through four coupled control planes: Expert Topology (what experts exist), Routing (which experts a token uses), Balance (how aggregate load is controlled), and Expert Parallelism (how selected computation runs on devices). The central claim is that the dominant trend has shifted from activating more sparse parameters toward decoupling semantic routing, computational budgets, and physical execution. If the framework holds, it gives designers a shared language for comparing MoE systems and for predicting which innovations relieve the next bottleneck.","feed_headline":"Four control planes explain MoE evolution, not release dates","feed_subtitle":"Sparse MoE's real shift: decoupling semantic routing, compute budgets, and physical execution.","key_machinery":"The load-bearing object is the four-control-plane closed loop: Topology defines the expert set, Routing selects an expert subset per token, Balance aggregates load statistics and feeds back loss, bias, capacity, placement, or replication, and Expert Parallelism maps the selection to dispatch, local GEMM, and combine on physical devices. The paper formalizes the loop as $\\mathcal{E}=T(\\theta_{\\mathrm{topo}})$, $\\mathcal{K}_t=R(x_t,\\mathcal{E};\\theta_{\\mathrm{route}},b)$, $(b,\\alpha,c)\\leftarrow C(\\{n_i\\},\\pi)$, and $y_t=\\mathrm{EP}(x_t,\\mathcal{K}_t,\\pi,\\sigma)$, showing that the four layers are not independent modules but a closed system. The companion object is the eight-milestone dependency graph, which records which bottleneck each historical change removed and what bottleneck it exposed, so the control planes and the milestone graph serve as complementary temporal and structural views.","core_discovery":"The paper's central claim is a re-description of MoE history and structure. Chronologically, it proposes eight architectural milestones tied to bottleneck migration: statistical division of labor, sparse Top-k conditional computation, Transformer-scale expert parallelism, open decoder-only MoE, fine-grained and shared experts, ultra-sparse scaling, token-dependent compute budgets, and semantic routing decoupled from physical execution. These are presented as a dependency graph, with nodes 1 through 6 forming a mainline and nodes 7 and 8 as stackable orthogonal branches, not as eight successive generations. Structurally, the paper claims that every MoE system is a closed loop of four control planes, Topology, Routing, Balance, and Expert Parallelism, with forward execution from topology to routing to execution and feedback from load statistics and system costs back to routing and placement. The paper's principal conclusion is that the main trend across this history is a shift from activating more sparse parameters toward decoupling semantic routing, computational budgets, and physical execution.","pith_inferences":["If the framework holds, the same four-plane decomposition could be used to compare MoE designs outside decoder-only pretraining, such as code models or multimodal systems, and to identify which control-plane knob produces a reported gain.","A testable extension is to adopt the paper's proposed equal-budget ablation matrix, fixed data, tokens, and active FLOPs with granularity, shared path, balance scope, and execution topology varied independently, as a standard protocol for architecture papers.","The claimed decoupling trend suggests that load balancing will migrate further from training-time losses toward runtime expert placement and replication, making routing stability under domain shift a more central research problem than new router families.","If the bottleneck-migration story is correct, future innovations should be sought where current bottlenecks concentrate: exposed communication, small-GEMM fragmentation, expert-weight I/O, and tail-latency variance under dynamic compute."],"forward_implications":["Comparing MoE systems by release date or benchmark ranking becomes less informative than locating them on the four control planes and tracing their bottleneck migration.","The modern mainline, token-choice Top-k routing with fine-grained experts, optional shared path, global or ALF balance, topology-limited routing, and runtime expert placement, is predicted to persist while frontier branches evolve as combinations rather than replacements.","Dynamic compute and semantic-physical decoupling are stackable with fine-grained and ultra-sparse structures, so there is no single next-generation architecture after ScMoE.","MoE evaluation should report active budget, exposed communication, specialization metrics, and tail latency, not just validation loss, because average training FLOPs do not guarantee systems efficiency.","Load balancing should be decomposed into expert, device, node, and communication levels, with runtime placement and replication treated as part of the balance plane rather than as post-hoc engineering."],"supporting_citations":[{"why":"Establishes the capacity lever: sparse Top-k routing decouples total parameters from per-token compute.","marker":"[5]"},{"why":"Fixes the Transformer plus Expert Parallel execution chain, Router, All-to-All dispatch, expert FFN, combine, that later milestones inherit.","marker":"[6]"},{"why":"Introduces fine-grained experts with shared-expert isolation, the granularity and shared-path axis of milestone 5.","marker":"[12]"},{"why":"Adds device-limited routing and device, node, and communication-level balance, showing logical expert balance differs from physical balance.","marker":"[13]"},{"why":"Supplies the non-gradient-bias ALF mechanism that replaces or weakens auxiliary balance losses.","marker":"[29]"},{"why":"Provides the modern mainline example of batch-wise ALF with a weak sequence-wise loss and runtime redundant expert placement.","marker":"[30]"},{"why":"Introduces shortcut-connected expert parallelism, the canonical semantic-physical decoupling branch.","marker":"[20]"},{"why":"Reports ultra-sparse scaling gains at a fixed active budget, the empirical basis for milestone 6.","marker":"[17]"}],"fun_headline_variants":["MoE evolution: from sparse activation to decoupled execution","Eight milestones, not eight generations: MoE's dependency graph","Four control planes unify MoE's design space","MoE's shift: decoupling routing, compute, and physical execution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework stands on the accuracy of third-party technical reports the author did not independently verify, so a misdescribed architecture or reported scaling number would loosen the empirical support for the milestone pattern and the claimed macro-trend.","fun_headline_variants_meta":{"raw":{"variants":["MoE evolution: from sparse activation to decoupled execution","Eight milestones, not eight generations: MoE's dependency graph","Four control planes unify MoE's design space","MoE's shift: decoupling routing, compute, and physical execution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000317,"raw_usage":{"total_tokens":1807,"prompt_tokens":976,"completion_tokens":831,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":761}},"tokens_in":592,"tokens_out":831,"duration_ms":8358,"temperature":1.0,"reasoning_tokens":761,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:27:25.157753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled equal-budget pretraining sweep, same data, token budget, and active parameters, varying expert granularity, shared path, balance scope, and execution topology independently, and measure validation loss, specialization, exposed communication, and tail latency; if no control-plane knob changes the outcomes, the bottleneck-migration story would be refuted.","supporting_citations":[],"review_version":1}