Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.715125Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.17644.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.715125Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
15 of 15 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3e40bfb8-68d0-4162-a39f-3ee33752e764 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda9b7c0-975e-49a7-a44d-c79672a8342d · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0b9cf3-8307-4248-9806-c49d1a1b584e · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f5b4d38-2a5a-4c5b-9986-7f2b0dbe516a · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Tutel: Adaptive Mixture-of-Experts at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc5e0e5-a6ad-4b3e-9656-cdfaab448c1c · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Accelerating Distributed MoE Training and Inference with Lina
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938b3a12-e031-4c54-bedd-6cc931241390 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2967adc6-a89e-4b02-aa21-be8e8d1c471d · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6107027a-ed16-41c6-ab09-fbc260fc3daa · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d53bd1c2-bea0-4529-b3e3-69c6ba285482 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepEP: an efficient expert- parallel communication library
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb99d2a5-fb3b-44ce-9700-19eed4d5b7f0 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7587e05-073a-4771-81f7-355e0ac466bf · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepSeek-V3 Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48fc9efd-7438-4d22-9406-373f497c5deb · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Striped Attention: Faster Ring Attention for Causal Transformers
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7d866cb-9aaf-4854-b798-14f33450d906 · outbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix arXiv:2603.02188
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.