Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:58.446447Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.08650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:58.446447Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 155412e3-7d0f-44fa-acc7-9d39d244d27b · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f7f081eb-159f-4567-ba8b-42f519014d53 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 88128731-e303-4cc8-b293-7508c5abf31e · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism 总 All-to-All 时间
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1b015d02-473e-477d-8c0c-00fff5e8c6e9 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78903f03-9259-4bf6-9fd4-697cb33ce04b · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Mixture of Experts in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2b9ad3-4bc8-440b-b148-5692f207abeb · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism A Survey on Inference Optimization Techniques for Mixture of Experts Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c66eaf-ae8f-4359-9f8e-fc8da2449351 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8984cb4b-f455-4def-9e56-094faf9ff28f · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Adaptive Mixtures of Local Experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a1ecfed-c799-4c62-858a-1ffa1c4724dc · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8cec16-c0d2-493a-953e-44a51c39c2ad · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c43a1dfd-0197-4799-aa11-1113006dd440 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c45be21f-5096-447c-9f5f-f1fe7941a958 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a98cf91f-94b7-4bdc-afd9-d0719badc588 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism BASE Layers: Simplifying Training of Large, Sparse Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32714d3f-e41a-46c4-90d8-7886c02c8b33 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture-of-Experts with Expert Choice Routing
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61850748-812c-48c4-941f-dec0ffaa1b60 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixtral of Experts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31688b3b-bc2c-40b5-9c15-8cc59de059d1 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ebc0aa-91b1-4fa2-a626-cbea7ee04639 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b20fc12-e4e1-4461-94ac-0d948153fc82 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Arctic: Snowflake’s Open-Source LLM
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 192b8514-451f-4f73-8613-4e8da2c1d24a · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Qwen3 Technical Report
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879cd6f2-c6d0-4b73-a76a-4efda08f2733 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Mixture of A Million Experts
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f73fa79-2c97-4189-945a-659621f98fed · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Kimi K2: Open Agentic Intelligence
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab426b07-485c-4686-a67a-560098855d76 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism gpt-oss Model Card: Architecture
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d634a2cc-ec24-42c7-8c8e-5debaca623cb · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat-Flash Technical Report
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6f590d-e17a-4d7c-ab6a-4bd47c855fca · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f27f14c-fbd6-44e2-998d-1f1b0f54fc92 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism LongCat 2.0 Technical Blog
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 898a66da-b382-4426-9b70-927251efbefc · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MoHGE: Mixture of Heterogeneous Grouped Experts
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6226d18c-7617-47dd-86bf-e488410062da · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism GMoE: Global Mixture-of-Experts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 42e1fceb-5840-425e-a95c-c0374e5b5b72 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Multi-Head LatentMoE and Head Parallel
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6ea333fd-a000-4440-a917-4a2dd234c26b · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7fe1c8-47ff-462b-8b70-68986a7d781b · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6bd2c283-d122-4b69-83c7-0fa779d27a17 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism From Sparse to Soft Mixtures of Experts
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa229c4d-40aa-428a-80f9-51b3769444ac · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d413417b-2018-4afc-b784-c10df13c43d0 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f6a2cc-cb6e-41bf-a386-d66cf431e1e2 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism DeepSeek-V3 Technical Report
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4dd52a-1fc9-48fe-891b-a29c91793666 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e0c6d9-4338-48fe-8537-2c9497c55301 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Tutel: Adaptive Mixture-of-Experts at Scale
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96585586-e007-4b53-a86a-6917574ca7c4 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism JetMoE: Reaching Llama2 Performance with 0.1M Dollars
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a825bfe9-00c2-4a9d-aa07-a6cb2e9774a0 · outbound
The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism Jamba: A Hybrid Transformer-Mamba Language Model
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.