Pith. sign in

Paper Citation Record · LEDGER

Sequence Parallelism: Long Sequence Training from System Perspective

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2105.13120.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.13120 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:10:00.348204Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.806179Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7f674bf-6ba0-432b-ad68-b49b25746a3f · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.164904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:0f1a0b7d82aed2c7232b6e79971bf144b321abe7b24b21c64209e87e7b15880b

Observation 03e34ab8-7ec6-42be-a8fa-b3947023d55f · inbound

World Model on Million-Length Video And Language With Blockwise RingAttention cites this paper.

World Model on Million-Length Video And Language With Blockwise RingAttention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:36:57.307966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T06:36:57.165551Z digest=sha256:d5e284023fc8a77d80e188c3fda93b6080f015bd60f9a5c9f107574ed0655fbd

Observation 664367ad-fd64-46f2-95fe-9a4d17909af8 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Sequence Parallelism: Long Sequence Training from System Perspective

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:27.874603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:bc8eb11d1e9934d7dc97f57b0d0b86947f0eaf1f6205fe411b26801b67c034cb

Observation 85eae24b-9be0-41e4-9ced-6facda989775 · inbound

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models cites this paper.

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T20:10:00.348204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:10:00.348204Z digest=sha256:e52e7d6b4833ad806de98270fc26e9643e3ccd45a73b19b536c68dcd6d361824

Observation 89950387-b391-4274-89af-6762633be046 · inbound

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training cites this paper.

When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training Sequence Parallelism: Long Sequence Training from System Perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:30:21.231024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:30:21.231024Z digest=sha256:5f94ffe610820f654cedb5db299e70c81e34a0f444c2dfd5f9b242ec3cc31adc

Observation 3a937a68-e871-4302-8c80-6e19c7abce6f · inbound

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries cites this paper.

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries Sequence Parallelism: Long Sequence Training from System Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:04:47.556883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:04:47.556883Z digest=sha256:f014b55d483b0b685487afb2aeacf8a1a7c8020d8d60d1f17d1fbca7d88fcf02

Observation 8458eaa1-9adc-4e5f-8467-ae62d3015b5d · inbound

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication cites this paper.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Sequence Parallelism: Long Sequence Training from System Perspective

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.920005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.920005Z digest=sha256:13185eeeedea38bf3cb33314bfd310466e81bfe805a5ebeb0d0aee7ba0b34a68

Observation 253fd9db-d347-4c64-ad4f-2a0c238ffe64 · inbound

Automatically Planning Optimal Parallel Strategy for Large Language Models cites this paper.

Automatically Planning Optimal Parallel Strategy for Large Language Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:59:59.133758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:59:59.133758Z digest=sha256:762d0847cb92c4983f4a497a89ba3d0f3f7b0dbcd09f7da04770362b6e27bd4e

Observation 9e654174-9035-441a-afb5-2289664958c6 · inbound

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation cites this paper.

HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation Sequence Parallelism: Long Sequence Training from System Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T21:16:16.508011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:16:16.508011Z digest=sha256:a2dbda019616166c1484a35ebcaa0effd59f6c0f4375827c087c709c16ddbf2f

Observation 2745746e-9eee-48bf-99c8-1b28bb474655 · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Sequence Parallelism: Long Sequence Training from System Perspective

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.343801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.343801Z digest=sha256:fb30a24f268c82a753ea3525101083f62dac3bd47228c3ddacbbcc348689a956

Observation 1c57709c-5045-4694-aebb-90fad6e68c7b · inbound

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning cites this paper.

Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning Sequence Parallelism: Long Sequence Training from System Perspective

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:55.653397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:55.653397Z digest=sha256:277b60384a54588b2db9948959dbfb7dab61fbeb72e883eef8dc935d0c589ec5

Observation 519d403c-33f9-416e-a771-6143d673a16c · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.986399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.986399Z digest=sha256:5ceb96f215e5f93f947b37747202f17cbf39bc0ca5133c5d00903fc9b9db3a0d

Observation 7b19a446-827a-4ff6-87aa-f8fe4cbc08b9 · inbound

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? cites this paper.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? Sequence Parallelism: Long Sequence Training from System Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:25.288961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:25.288961Z digest=sha256:cffbf3ed14ffac363ff59da2cfa2d2b6849ad3873eba0b3151ad722f18600ec2

Observation 8316e377-e9ca-4bba-99fc-99db6a301f91 · inbound

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs cites this paper.

StreamBP: Memory-Efficient Exact Backpropagation for Long Sequence Training of LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:36.333821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:36.333821Z digest=sha256:c006280ad476ba2c16b5a93287f54de275e884556f54a820938ff68cea104278

Observation aa75cc33-de9a-4262-beba-b1b0be105839 · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute Sequence Parallelism: Long Sequence Training from System Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:32.085563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:32.085563Z digest=sha256:aa9c33d76eec74d84746e5f8c398b81d43118495ba17fbf26cd77ce315a445fa

Observation 4397b7e9-5a51-4462-97ae-4f0f935b4ab0 · inbound

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library cites this paper.

Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Sequence Parallelism: Long Sequence Training from System Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:32.922834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:32.922834Z digest=sha256:a11a7bd9ae9ca0994f7dd848d62159f91c141302843f2f9900b18dac3158c6a4

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:37682624de5a83e7949403da4e3f8c00e65fddb9ce59d8152ff4c7a5b510f0b6

Observation b5dd0957-cac0-43ba-8316-4e8de9778731 · inbound

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs cites this paper.

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:08.905826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:08.905826Z digest=sha256:76db220b1dbecaedec8471e6ffe9ba94fdd27b336aecad62fa83d330ac00c4dc

Observation dc945350-74c7-40d7-a7fe-22216310527c · inbound

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference cites this paper.

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T17:55:11.234042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:55:11.234042Z digest=sha256:b25bc64a0ddd5760a07faa09eaf5082bf99339fdd5075a1c8acbf8beaeaff83f

Observation 5b76fa2b-c631-4d1f-91e3-164d2e7593a7 · inbound

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference cites this paper.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.794986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.794986Z digest=sha256:3ded3e61ada9212b91e448d523f880fcacadd546352d6a50c49cd45673daf2d2

Observation 547516d7-5037-43eb-a675-17713898ed44 · inbound

TetriServe: Efficiently Serving Mixed DiT Workloads cites this paper.

TetriServe: Efficiently Serving Mixed DiT Workloads Sequence Parallelism: Long Sequence Training from System Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:10.837066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:10.837066Z digest=sha256:2a3711c94f65bb20ad8266ad5985368d879ebd8ab0a303ae135db4b0275a9d5a

Observation 0a064b01-f126-40e6-bf92-a8c2d30f1b24 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence Parallelism: Long Sequence Training from System Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:50.032377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:50.032377Z digest=sha256:3f51181f317b0abb9ea7700b7a5c7b43ef72e94dbdc931235e477b790df6e2e0

Observation 02222285-4611-4fc1-8a99-ed855fe0cc04 · inbound

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems cites this paper.

Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:14.135198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T04:41:03.457680Z digest=sha256:bcdb7c15aae27ec082cce63568d19117704e51b5b79c030562e2fdc5f249c335

Observation a43b33c6-db05-495b-975a-439585037437 · inbound

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference cites this paper.

Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.770043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T08:55:31.298030Z digest=sha256:00683778330ce961c968492acc245018060d38289ca1bac5eb93b12c9c71fe58

Observation 442c23a1-c1ce-4b8c-b020-79353d910f77 · inbound

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing cites this paper.

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing Sequence Parallelism: Long Sequence Training from System Perspective

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:53:24.750202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T07:53:11.761843Z digest=sha256:c49a77fc32ce3b69fbb023cbc9904c851207394fefb78338a2d30b61fc09f239

Observation 9525b37c-6ba3-41ad-81f2-6b08c1687abc · inbound

Online Dynamic Batching with Formal Guarantees for LLM Training cites this paper.

Online Dynamic Batching with Formal Guarantees for LLM Training Sequence Parallelism: Long Sequence Training from System Perspective

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:29:35.935578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T15:57:18.544274Z digest=sha256:12d19ae06108cabf2d7431729746a2e4b71740c9c53a0cf38d378de7064f2712

Observation a02cb5f7-b5af-431a-a4ad-2b877e49476e · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill Sequence Parallelism: Long Sequence Training from System Perspective

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.807624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:8d56ed9fb7413e7c4473d20191b41bf2f96b3729b2dad6fc1391a8fb52b63ac8

Observation 69a5c0ea-6441-4a69-8384-16c246a3f958 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:28:43.956670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T17:26:07.870260Z digest=sha256:4977e08c0fea1246275221b5e92fab0fa74d2012ca32b9680f009819d8485abe

Observation a648ba3b-41c3-4a22-b4a4-38b97f72cd32 · inbound

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint cites this paper.

PHOENIX: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint Sequence Parallelism: Long Sequence Training from System Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T08:40:29.554554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:40:29.554554Z digest=sha256:3e32250ef5f42d9176fefce7a0052df792bd5840a740c021ed2e7afde8b8458d

Observation f8b4783f-dd2b-4cb4-8f77-fd28935a3bad · inbound

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention cites this paper.

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention Sequence Parallelism: Long Sequence Training from System Perspective

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:27:41.863772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T06:22:23.985308Z digest=sha256:d80609603022fc7d08501a8839c9629a79180518158ce6d25a41a341b6fdb31c

Observation 6712b88c-b529-44d7-87b3-51e4c5d5f696 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Sequence Parallelism: Long Sequence Training from System Perspective

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:f8a636ba69c9b659fabf912077b8acc5fb53e078e2b18ce51184a33013b46543

Observation ad4c00d3-51d0-48fa-ab5b-91bdf8133d7a · inbound

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving cites this paper.

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving Sequence Parallelism: Long Sequence Training from System Perspective

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:14.069087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:54:14.069087Z digest=sha256:1a86714f2c18a947db003a3ad1977a9108c6e5426eee7b0f09b39fd6594acf5d