Pith. sign in

Paper Citation Record · LEDGER

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference

As of 14 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.00217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00217 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:52:30.795158Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24514ff3-c30a-4254-a4ec-ed5a690fdb0a · outbound

This paper cites Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.348753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.348753Z digest=sha256:c9acf5d0860c93808c7635a77fde4f8d6e53fcb2395a04d139f210835df2dd96

Observation 6a79779b-d657-4fdc-a734-2e37e0500b3e · outbound

This paper cites Beyond data and model parallelism for deep neural networks.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Beyond data and model parallelism for deep neural networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:33.107795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:28.503168Z digest=sha256:ef35f0e6a149e2fc5de592f9d40b9ee15b4fb657179e866db9f465ed775079a5

Observation 40a14d2e-66cf-4780-9500-d1b0f6765978 · outbound

This paper cites DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.621920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.621920Z digest=sha256:09077a89f42a11430caf93173db6c5f89e9e5c327278302fc15ced4b443d21c5

Observation 5b76fa2b-c631-4d1f-91e3-164d2e7593a7 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.794986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.794986Z digest=sha256:3ded3e61ada9212b91e448d523f880fcacadd546352d6a50c49cd45673daf2d2

Observation a7f1be0c-89e7-48d9-b685-2dde9885f28f · outbound

This paper cites Uniap: Unifying inter-and intra-layer automatic parallelism by mixed integer quadratic programming.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Uniap: Unifying inter-and intra-layer automatic parallelism by mixed integer quadratic programming

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.938616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:28.980821Z digest=sha256:e8eb244c2aa49f0962bc0258c365b54f49f5db379646112d279bd7b304160f7a

Observation 7b101fe8-429b-4831-b6ba-58386779095b · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.128897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.128897Z digest=sha256:6814d525a12e9c6d71293657b605ea5d4913991b789572151fea4a790a34ee3d

Observation 1f5b5377-1a15-4f07-84e9-84a7601fa0a8 · outbound

This paper cites GTC 2024 Presentation Slides.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference GTC 2024 Presentation Slides

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.719159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.248129Z digest=sha256:f8e13ee82dc1413312179f6923662a2edd46acb0958cdb23e40246679b3cebec

Observation 2228811a-d1e5-4ec9-ba21-6477f006a408 · outbound

This paper cites XLA: A Machine Learning Compiler.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference XLA: A Machine Learning Compiler

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.537233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.393837Z digest=sha256:64db70e9bca5566d7ea389a80382d82fabfae21b7a457efc13e10e72cbaa6770

Observation 595358c9-665f-4332-8376-9a535c1632b6 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Pytorch: An imperative style, high-performance deep learning library

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.531030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.531030Z digest=sha256:55f1cbba44d4b94a441eba259e9d030c6600bf5eeccdb6ed0c57c6855cb038d1

Observation b95ac045-07ad-4f81-933b-504736ffec3a · outbound

This paper cites Chimera: Communication fusion for hybrid parallelism in large language models.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Chimera: Communication fusion for hybrid parallelism in large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.324083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.660202Z digest=sha256:fe67319a831e3ebb9abb7e8ae2323d18687d5297b83fc957059b8c12ccdf59a1

Observation 8807b12b-2a57-4ff6-86fc-7823c750932f · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Stable-baselines3: Reliable reinforcement learning implementations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.123443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.769958Z digest=sha256:e405e7b1093e8748732adeb51bcce978672de6065ef513e6f287bfda6ec67b7b

Observation f477b404-937b-412b-9dc5-083a2c8a6dc7 · outbound

This paper cites COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.883265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.883265Z digest=sha256:e4482e755a83c203f0b9fd618139e33bfa9d4fd7914cdc586a2cd7e3e1aec392

Observation 8c328cf8-f322-402f-ac26-24a7ec6b5976 · outbound

This paper cites TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.961252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.961252Z digest=sha256:3460971f374642cfa657f9990426bf091f1974eb21d6b58f078afd5e228902b9

Observation ad31324f-7da3-4822-8c0c-b32818347c4b · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Megatron-lm: Training multi-billion parameter language models using model parallelism

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.925984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.105334Z digest=sha256:31e460aad7f864b814b808d4009e640c95158237e6e3ac8e52f24d2219c2777a

Observation f90f5743-a83d-4467-9673-c90090343477 · outbound

This paper cites Seesaw: High-throughput LLM Inference via Model Re-sharding.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Seesaw: High-throughput LLM Inference via Model Re-sharding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.186309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.186309Z digest=sha256:a8852b4ad7da03910f3ca18e11c89d61415c4551ab66f1822eb8129993983119

Observation 71fae5b5-2e42-40e4-ba1c-1b99cd58f76c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.295978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.295978Z digest=sha256:76e2f2fc7b7230fbfd408a8fe37b85580e4168c0502909b4e1832c9862fce858

Observation 1f22e790-a78a-4de0-b34c-f9ac21ef6a9b · outbound

This paper cites Gspmd: General and scalable parallelization for ml computation graphs.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Gspmd: General and scalable parallelization for ml computation graphs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.726728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.383478Z digest=sha256:10323c232cadf094d63b646242453a5898a7f413b931b5ff244618aeddacb45c

Observation a7e2a1a5-b25a-4ed8-a26f-354291a9a81e · outbound

This paper cites E., and Stoica, I.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference E., and Stoica, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.519861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.482618Z digest=sha256:d732969e431e5c3d83bd11ba40fed9cc1a1ecc33dd7994800dfc2e30c6b04802

Observation 2db2af95-05f3-4256-babb-53c556d30641 · outbound

This paper cites E., Stoica, I., Jin, X., Xing, E., Chen, Q.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference E., Stoica, I., Jin, X., Xing, E., Chen, Q

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.283647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.596691Z digest=sha256:0c2bb7fbce363c58b5a37e134875dfc4bdccc53e27d8929df9c9a34111deb029

Observation adf31cff-bb2c-4e4e-8fdd-62fbfd84edf5 · outbound

This paper cites Transferable graph optimizers for ml compilers.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Transferable graph optimizers for ml compilers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.107411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.710769Z digest=sha256:4010ef446e591ffdff05a6cafe5a92927c570b167436e1eacc404c7b9c044285

Observation 4cbc4e50-ff62-422c-bcf1-4b136548b95f · outbound

This paper cites write newline.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference write newline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.795158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.795158Z digest=sha256:b8ad32b05d3c7f44c058ae3c2116dcd318a3ccaf4eac8bcd4264644946b0f16d

Pith citing papers

No inbound Pith citation observations are available.