Pith. sign in

Paper Citation Record · LEDGER

ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.21231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.21231 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:41:45.459582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:27:28.048568Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8017b579-2402-4f6b-95f9-c90fcc2523bd · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.582899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:9588ca0a9947e914bbb40046f14e68f56a6cd7b33e081a897e53009384fd5baf

Observation d062ed69-fb33-4f4d-9da4-7b47739030ca · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:31:15.775808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:1a98fd7e6121d02e4f5db2d03754f8e49daee061818f5dca322d33d9d7534694

Observation 033a95d0-8d15-4b90-991a-67655237a820 · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.261784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:bb57ba074a8e3528040c700ecff0d1a1e6f89056988173f2ea04c74e481a7815

Observation a492123d-5b69-401e-b410-6bbfaff96408 · inbound

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training cites this paper.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.608821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:0bfef385e4941db42e35641ecce629760398cfbbe0d7efe9ba2e955aa55525fb

Observation dc9975fe-da9b-44d6-a3bd-52eb37215dba · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:40.953528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:c6edd117ea2e81676118d27e2efe8186a4058b7125e8b70a4a3980e4898cfbb3

Observation e5ff70bb-1d1b-4c95-87ef-838c8574ecd0 · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.395862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:ed6b2aa9116e31caf9e55106d494dc8c3a85f88cb26a8b065cc15b30c1187c5a

Observation c1b504b3-a5c5-42c3-8af3-6e14e81c1f74 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.982178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:847793a1b3127d7b7ff38ef9558b467e69c25b4813dc7883096784b78d6fff1a

Observation e25cb165-2ed7-4dd4-b56c-c6e1d6cf8167 · inbound

FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training cites this paper.

FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:27:28.050061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:10:34.962717Z digest=sha256:a16d5c13dcb27841307465b779b2d3285e4e1bf7b85a87773a9cc1e66a4bb2bd

Observation 4e602654-14e9-497d-ab0b-7d5d49199b30 · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.459582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.459582Z digest=sha256:fbf189328b69a6479cfd5eb777eb53799f271f4b5801ebe1263d63ceaecd9f19