Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:51.292724Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2602.21196.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T21:11:51.292724Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c99560fe-254d-461e-b429-c9517ab8539e · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b7a624e-99ee-42be-84d2-49e96250f034 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68b75ca-1d76-4926-b339-af583cf09ac8 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afd8f254-9ddd-4193-9ac6-e7ee3efeba62 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96885c7e-b90e-4767-8fde-157aae7cddaa · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e2d119-88ba-4606-bd02-3953d6af281a · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Y., Ermon, S., Rudra, A., and R \'e , C
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cda7e97-6288-4297-9187-bcb71248e09d · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f768da-03cd-436e-8b4a-4fb7ae1a4679 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Gemini 3.0: A new era of intelligence with gemini 3
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55cd6d9f-3cba-42f5-8de9-4db081f938f5 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b8a64e6-f660-40a4-99ac-8949748bea35 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Unsloth, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dcfedae-0636-4195-a19d-59acb3f1eb3f · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Advanced Long-context End-to-end Speech Recognition Using Context-expanded Transformers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4190dcf3-bdb3-4e09-a596-5bd0f0cdfc3c · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Liger-kernel: Efficient triton kernels for LLM training
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001bd1f1-5071-4097-8ae6-d3ae4e2eaf8c · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Qwen2.5-Coder Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea053d39-524b-43bb-b29e-3471a748a529 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96becb5-636f-4967-96d9-4265da21b165 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Mistral 7B
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25749982-9c90-41c1-a6d2-5bdb6013d582 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8edc2120-d570-4a4a-8b49-e52572cbe760 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Kimi K2: Open Agentic Intelligence
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5563856c-5726-4244-98e2-77337331e676 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking StarCoder: may the source be with you!
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a064b01-f126-40e6-bf92-a8c2d30f1b24 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence Parallelism: Long Sequence Training from System Perspective
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67abdf03-4eaa-42b9-a064-545937294de6 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Sequence parallelism: Long sequence training from system perspective
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee07d3f-0a17-4640-afd9-3d4765fae34e · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Torchtitan: One-stop pytorch native solution for production ready LLM pretraining
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2082e951-a090-4af2-8b19-a5badeed4f86 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50b51a8-d40d-4473-ad95-6f265343e380 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking MAGI-1: Autoregressive Video Generation at Scale
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9875024-4b17-4f39-b112-889bd3585726 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5aa8ca-ed68-452b-8ff0-a1db0ab6c9c8 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ff2873-f83a-469e-ade1-6898833286c3 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Wan: Open and Advanced Large-Scale Video Generative Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f46242a-be0f-4f88-a349-e6cf3f157d53 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking N., Kaiser, L., and Polosukhin, I
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83960181-6e2b-4ede-8d1f-f2233a667b44 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking HunyuanVideo 1.5 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb410f2-0c61-40f6-b4b6-97a6acbd66e3 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78d2d07-028f-4631-9aff-1f837a23dda8 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7629dfe-3273-4b4e-95e1-5320c0571101 · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4cfafaf-795e-49b0-bafe-c0b38fc668cf · outbound
Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.