Pith. sign in

Paper Citation Record · LEDGER

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

As of 19 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2506.13996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13996 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:57.837571Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:48.095184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40367eda-1cf8-4460-8d22-17d7ec9af6c2 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.473062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.473062Z digest=sha256:14243799c72f3df039c9678f57b896729fa375d3572ef0339954fb63edf2931b

Observation 97957d92-9f4b-4298-bd7f-1b37b5bb3d3a · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.703220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.703220Z digest=sha256:e22e0516d0a3ea5f34bd017c079f762d4d2704ab3f5ed762a619ae00a57f63d6

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · outbound

This paper cites Reducing Activation Recomputation in Large Transformer Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:73cca6a51af1685ed7f389c896dfe3e86653e7cc2dde76564a7c71ade3e244a0

Observation 9786b61f-a944-441d-8011-a71f08b31f5f · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.105856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.105856Z digest=sha256:eaabb84546486f06d4124f376851537f93e2f8de58a3bc6204b5d6215287f58a

Observation c4684dc0-5dd9-423f-b97d-1dde205943c9 · outbound

This paper cites DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.230466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.230466Z digest=sha256:fee9ea0de13be69811659cd60e345e65dfa6a7538cbda7dd4545b376eb1b111c

Observation 4cf2c8ee-c34e-4f71-a054-6050612b2b61 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.281690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.281690Z digest=sha256:610fdd3c72b4cc323d02668d3fc9d67d0c9c54f2866fdb594d29ed85d1af079e

Observation 9d9e0298-ab80-4bb6-8aba-7eadca48f60d · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.362359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.362359Z digest=sha256:e8febd6ca6340900753ebf81d79cfcd66db4e37836af6c1af300170150e24bc1

Observation 3d33a8a5-0944-4486-9b2e-f980395fd95f · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.510416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.510416Z digest=sha256:85bb219b41bbd4104f6bff71679f99c79fa82dde34dc2fa170701e9611f1bede

Observation 56cbb9a8-b30c-4123-8fb3-a696124d50ce · outbound

This paper cites LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.602093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.602093Z digest=sha256:8143af441dbdba4930127f053814b749f77050a2437c394fbf0db31e8858882f

Observation 6ba52778-8e90-491f-a35d-e8bedcd57bde · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.663570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.663570Z digest=sha256:8833aee7da3126ee6c68833e617ef3cf728397bfbcc8f1ae2ad09ea431e9b84c

Observation dd1ff85e-5006-4493-ac35-44d4b275336f · outbound

This paper cites Datasets: A community library for natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Datasets: A community library for natural language processing,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.266013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:30:57.742706Z digest=sha256:07c0517c2243fae77320007da35f6762fe4cc5e4c17c57db32c3aeee7ba085a2

Observation fd638e83-7865-4f24-b9b6-1e00006d7cc6 · outbound

This paper cites ArcticTraining: Simplifying and accelerating post-training for large language models,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ArcticTraining: Simplifying and accelerating post-training for large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.250240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:30:57.748571Z digest=sha256:384894d36f52f7118a25010181512e4233409d085ccfda802d5e303dd4236579

Observation 622566c1-dd5c-448b-9a97-c4efac6fe13b · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.786155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.786155Z digest=sha256:83982c4f245363a077de5d6bd3df09dfd0be4ae4e1beb66d69be563f25258b0b

Observation ca5d832f-b92b-41ae-a63e-3db5732e2929 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Transformers: State-of-the-art natural language processing,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:30:58.233466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:30:57.805996Z digest=sha256:c921cdc0e9de51c833bb7776369b98e917b52099dcae5932d2606b750ab77d78

Observation 26203125-424d-4dc8-ac48-05c8314a0a3d · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.822249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.822249Z digest=sha256:1226ad0cba932a0bf72641000bb302091aad9b8d4de2d0bf60813321f7a5cb61

Observation 19ccebd1-f9c0-45b0-91fe-236b9f080d5e · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.832992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.832992Z digest=sha256:183450abf5a2037cc6b9dbc33e97d15ea59d79d7bd26ea19479eace4090171ed

Observation b8d73282-b716-4528-85b6-a516858600f3 · outbound

This paper cites The Case for Co-Designing Model Architectures with Hardware.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences The Case for Co-Designing Model Architectures with Hardware

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:57.837571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:57.837571Z digest=sha256:ae84de5a9f26d5cb8ed3d5424a0efaeb9b4d3117439de82fbcf0c14c04090a6f

Pith citing papers

Observation f68b75ca-1d76-4926-b339-af583cf09ac8 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:48.095184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:48.095184Z digest=sha256:4abba694c1f1e490e68dfbd2fcfaa61e8475e165eee6c66470b714e2fb5cb258