Pith. sign in

Paper Citation Record · LEDGER

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2607.17644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17644 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.715125Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e40bfb8-68d0-4162-a39f-3ee33752e764 · outbound

This paper cites Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:36.731755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:36.731755Z digest=sha256:a8a7cfd1670a5ad5421991b40b2f96a8cf814495b74a09fb44b3564ba50ed743

Observation cda9b7c0-975e-49a7-a44d-c79672a8342d · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.049832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.049832Z digest=sha256:c7ea393c0abc34c2cb681e8191958aff611657034d77b13d9718d7e554097426

Observation fb0b9cf3-8307-4248-9806-c49d1a1b584e · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.115161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.115161Z digest=sha256:5668df09b5307998b325e574dd1882d3afd544dda568c8392c1d9a26026f4f63

Observation 1f5b4d38-2a5a-4c5b-9986-7f2b0dbe516a · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Tutel: Adaptive Mixture-of-Experts at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.181249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.181249Z digest=sha256:174a9dd9ad36b8f58e3f5355e1439f53d8c0563f63baf5a621ed7eac9e14338e

Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.263658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.263658Z digest=sha256:1364f018d5018818a3088a48b333c67642793562a0312096fa204a1a0e01d6a4

Observation 6dc5e0e5-a6ad-4b3e-9656-cdfaab448c1c · outbound

This paper cites Accelerating Distributed MoE Training and Inference with Lina.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Accelerating Distributed MoE Training and Inference with Lina

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.332884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.332884Z digest=sha256:3ef196b06de72e046636e26fb99a1659a3b19214711ee658f9738cff16300915

Observation 938b3a12-e031-4c54-bedd-6cc931241390 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.392437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.392437Z digest=sha256:d1c8efc7b276bdf634001b27131a2ea1beeebe324831b570e7c83df0d752fa94

Observation 2967adc6-a89e-4b02-aa21-be8e8d1c471d · outbound

This paper cites LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.570769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.570769Z digest=sha256:95e276c343c9efb098fb095978fd26d8f1c162350a82570b7150b10ddc56b4d8

Observation 6107027a-ed16-41c6-ab09-fbc260fc3daa · outbound

This paper cites TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.642188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.642188Z digest=sha256:4772968c71fe8a357d7ea63796092c7694bfb194030ad45aadaaf02e90086983

Observation d53bd1c2-bea0-4529-b3e3-69c6ba285482 · outbound

This paper cites DeepEP: an efficient expert- parallel communication library.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepEP: an efficient expert- parallel communication library

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.715125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.715125Z digest=sha256:9fb6dfe0366afb4f2a8d4ad7d19dbe9ee62a7ff6e4c9c60693578af44e09cc32

Observation 73494bae-8a1d-4c9f-a9ac-3a3c3d90d1a0 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.511773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.511773Z digest=sha256:e2c2f80ede73c87b07dc8dc8373544bfcd54c33aed392bf1b1cbdfeaaca38055

Observation cb99d2a5-fb3b-44ce-9700-19eed4d5b7f0 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:36.906798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:36.906798Z digest=sha256:6a277b4a5481618aad8134b29b4a6440932839d3c7bf2d1c12ea7da62c858985

Observation f7587e05-073a-4771-81f7-355e0ac466bf · outbound

This paper cites DeepSeek-V3 Technical Report.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:36.988309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:36.988309Z digest=sha256:c25b5fa1182e96caa8a866fc1d593f0edd587ed986427f3b10ca13a9990d8a95

Observation 48fc9efd-7438-4d22-9406-373f497c5deb · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Striped Attention: Faster Ring Attention for Causal Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:36.816550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:36.816550Z digest=sha256:ec8dca3b993f1b4ce4dd3e17d5ae5aafb78d0695ffdcd9d666e5d8f9dea898f4

Observation c7d866cb-9aaf-4854-b798-14f33450d906 · outbound

This paper cites arXiv:2603.02188.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix arXiv:2603.02188

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.454051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.454051Z digest=sha256:1bc914f4f567c9429dcc9f593b5454e0c94a4d31f07959b582b8782c28eef448

Pith citing papers

No inbound Pith citation observations are available.