Pith. sign in

Paper Citation Record · LEDGER

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2509.00217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00217 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:52:30.795158Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24514ff3-c30a-4254-a4ec-ed5a690fdb0a · outbound

This paper cites Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.348753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.348753Z digest=sha256:1faa0ae42a1040d4b49db05168c38bc19baed48d0c32aabe759f1dc1a5c9ff48

Observation 6a79779b-d657-4fdc-a734-2e37e0500b3e · outbound

This paper cites Beyond data and model parallelism for deep neural networks.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Beyond data and model parallelism for deep neural networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:33.107795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:28.503168Z digest=sha256:2d79576557dd3aea9b987ef3722d66f4fb4a1657831587a0f96811629aa9ab16

Observation 40a14d2e-66cf-4780-9500-d1b0f6765978 · outbound

This paper cites DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.621920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.621920Z digest=sha256:e091f1bfed7670f381b13f8255974423bab6b821d7d65e31ddce6c7743b57bab

Observation 5b76fa2b-c631-4d1f-91e3-164d2e7593a7 · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Sequence Parallelism: Long Sequence Training from System Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:28.794986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:28.794986Z digest=sha256:2a0918bff38375e980a38bb87a56249704557abfe84973e59067a57c7e8324f5

Observation a7f1be0c-89e7-48d9-b685-2dde9885f28f · outbound

This paper cites Uniap: Unifying inter-and intra-layer automatic parallelism by mixed integer quadratic programming.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Uniap: Unifying inter-and intra-layer automatic parallelism by mixed integer quadratic programming

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.938616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:28.980821Z digest=sha256:ce0067fd9c76514c55146d7ee874b74ac56ba821179dc429234925e98ad5f990

Observation 7b101fe8-429b-4831-b6ba-58386779095b · outbound

This paper cites DeepSeek-V3 Technical Report.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.128897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.128897Z digest=sha256:8c54ac6324d67c1e70ab948718a2ca55c5b5ccde4e0a9fbc6893160f73008658

Observation 1f5b5377-1a15-4f07-84e9-84a7601fa0a8 · outbound

This paper cites GTC 2024 Presentation Slides.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference GTC 2024 Presentation Slides

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.719159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.248129Z digest=sha256:85769ec85c60b621495d65dc732de6415127b25d7a88496ae140d289ca7b0028

Observation 2228811a-d1e5-4ec9-ba21-6477f006a408 · outbound

This paper cites XLA: A Machine Learning Compiler.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference XLA: A Machine Learning Compiler

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.537233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.393837Z digest=sha256:7ebc1b0a50eb63a7f5a86964c01bb17d28def321574040b10da20621f046ff4c

Observation 595358c9-665f-4332-8376-9a535c1632b6 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Pytorch: An imperative style, high-performance deep learning library

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.531030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.531030Z digest=sha256:20241423f18775039a4970cdb782c688edec018461a7e8afca13c3083331c5bc

Observation b95ac045-07ad-4f81-933b-504736ffec3a · outbound

This paper cites Chimera: Communication fusion for hybrid parallelism in large language models.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Chimera: Communication fusion for hybrid parallelism in large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.324083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.660202Z digest=sha256:8448da48cceef0da673d7b8c6d135099d588c642c7f07e42aec9d54cb72d6f3d

Observation 8807b12b-2a57-4ff6-86fc-7823c750932f · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Stable-baselines3: Reliable reinforcement learning implementations

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:32.123443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:29.769958Z digest=sha256:716392a33fb726191b58390eec9367e4a5f558c887dedfe69770f43cb1e5ed9d

Observation f477b404-937b-412b-9dc5-083a2c8a6dc7 · outbound

This paper cites COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.883265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.883265Z digest=sha256:76edb52b8fd988c7a8c9e3d382c90115c6e1a1350075f5efec8154f64d559702

Observation 8c328cf8-f322-402f-ac26-24a7ec6b5976 · outbound

This paper cites TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:29.961252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:29.961252Z digest=sha256:30d949d87c9fd89abf88ba25e8bbb5bb2c65c5d024bc815e6b93fe8ce8aadf1a

Observation ad31324f-7da3-4822-8c0c-b32818347c4b · outbound

This paper cites Megatron-lm: Training multi-billion parameter language models using model parallelism.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Megatron-lm: Training multi-billion parameter language models using model parallelism

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.925984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.105334Z digest=sha256:37d6e426c12b7f5b0426eee71ab0088f6f4b0ed740c8da23e355516a6b70ff44

Observation f90f5743-a83d-4467-9673-c90090343477 · outbound

This paper cites Seesaw: High-throughput LLM Inference via Model Re-sharding.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Seesaw: High-throughput LLM Inference via Model Re-sharding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.186309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.186309Z digest=sha256:53159ea24a763a7758c76b52e80bb29afc5a4bfb9caf21d850e3ff3b4689b83a

Observation 71fae5b5-2e42-40e4-ba1c-1b99cd58f76c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.295978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.295978Z digest=sha256:c8a6ee885b3e8857dee6b96b1cbaedf3aadd4db978b7520cde982ad1015c129e

Observation 1f22e790-a78a-4de0-b34c-f9ac21ef6a9b · outbound

This paper cites Gspmd: General and scalable parallelization for ml computation graphs.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Gspmd: General and scalable parallelization for ml computation graphs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.726728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.383478Z digest=sha256:c0ec2ab5239496427fdee38b4765c3f2970879ba7d6b5b153af6cc8eecf36b50

Observation a7e2a1a5-b25a-4ed8-a26f-354291a9a81e · outbound

This paper cites E., and Stoica, I.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference E., and Stoica, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.519861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.482618Z digest=sha256:ad9c9b72f1ac3f226336a4fc559c2112e00a44e47a77cb02ad4cd7bea495c4ed

Observation 2db2af95-05f3-4256-babb-53c556d30641 · outbound

This paper cites E., Stoica, I., Jin, X., Xing, E., Chen, Q.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference E., Stoica, I., Jin, X., Xing, E., Chen, Q

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.283647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.596691Z digest=sha256:8acb7cde142af0b74c633323c66e90e4c10e292235c174ccd62ec8fd2475f9c4

Observation adf31cff-bb2c-4e4e-8fdd-62fbfd84edf5 · outbound

This paper cites Transferable graph optimizers for ml compilers.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference Transferable graph optimizers for ml compilers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:52:31.107411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T13:52:30.710769Z digest=sha256:9e4b94905da18e255154e1611019c0f45e6e6b39b1824b34de734eb6408e375c

Observation 4cbc4e50-ff62-422c-bcf1-4b136548b95f · outbound

This paper cites write newline.

Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference write newline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:52:30.795158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:52:30.795158Z digest=sha256:7617b6770a5557f3f8404986920f28b62c7a759201dd731274fde6b629c2ff97

Pith citing papers

No inbound Pith citation observations are available.