Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2207.00032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2207.00032 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:18:07.705592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:59:03.322421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39900118-976a-4971-8a9c-d9c4cbe878d9 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.431176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:10a678338d86240fffe1e8fc5d367cca152dd1247d0e6886457defc2fe6d20a9

Observation 08dccdd5-98d0-49e9-8fc2-68424e18710b · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.819111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:ae03a323ed793b45e6877e9693241b53837bdf649e795a7045c9f7de8ac7e9ec

Observation dd5d2662-76b1-44e0-91a3-18f9d06fb1ef · inbound

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making cites this paper.

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:18:07.705592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:18:07.705592Z digest=sha256:30939770e8a7b59ca210209522ec006d7cc4b745a56fe6870f7e4d842d19a04a

Observation 58157a7a-196f-4952-a0fe-cc8c9c97bf12 · inbound

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation cites this paper.

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:16.021201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:16.021201Z digest=sha256:de2e282b20750a4926f056e3fb0f91e71a47a99295f73d3e0c23714affa8155b

Observation c8abc7c3-bda3-4139-8f81-e0cb39023140 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.198408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.198408Z digest=sha256:d10a7710546117f77e0f530a6aa3ff8d88ebb8b95355a5504b3153a0a42ecf5d

Observation 2f5b486c-3f15-42ee-8066-5e3bcf0cabab · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.546676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.546676Z digest=sha256:45b1fcf9b1968f34d5d0e57ea1d18694d9565c2861f39f6e1b9f026efbe03e47

Observation 1bc7a03b-05fc-4f7c-abe1-632d5628a92a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.268813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:4c221df0d700227237d924dc239807b81088e48ed7f30cadd0975090d16dd99d

Observation de56e279-cca7-41aa-b6f5-8e0a892fad08 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:29.589719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:4193c59d8113b9232163a2754d05746d8836a5e5629873c89a8a09ab562aeb41

Observation 48682b07-9fa5-4f44-9259-f8e962a6d74b · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.439048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:8ca91a00cfecfdd37bd36cec6c7b5bd99c4df0b1450108435ebccbf9b9f4d915

Observation 97c1c357-ece0-4476-ad31-83fd25f2782d · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.897769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:18:05.789417Z digest=sha256:b1ca8e71a396fdf3cb54ddd8d45562b5f35bd3c092e2391984a0cf12edc2acc1

Observation 4b7cea5c-8342-4858-b4d1-0379d2ba33f4 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:49:48.919062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T06:49:16.271699Z digest=sha256:ca80099742d85ada4f3d3be9119722972fc1b19b01095367aaeaa5dc2011eca8

Observation e122ca3f-c6ac-4e43-8450-0c2acab501e7 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:43.898162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:43.898162Z digest=sha256:a78f02b231a938744694a74e3efa4724246c5611d3c45057e0e7dc83b8b614cd

Observation bafbf0de-24ce-459a-8c15-d717f12901b7 · inbound

ShardTensor: Domain Parallelism for Scientific Machine Learning cites this paper.

ShardTensor: Domain Parallelism for Scientific Machine Learning DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:09.104038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T03:03:32.316700Z digest=sha256:77500f6591258885e3962c888baca8bcaf8eb14a4a3d9a316521418c763c3aca

Observation a83bf6aa-9455-413c-b910-edc84e1334a9 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:5fe56c9f79fb9431ae18752bff6bc2ea5b78d164018a9b88cdedded50e24cb12

Observation 0bf0fca9-42f7-4096-af22-a01062156484 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.843565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:6ba2fa3879fba70475f4121a5918e244576edbe17d1c316e819871fcde2686b4

Observation 2ca665c7-8363-4c6b-8dd6-93f5aa9b6272 · inbound

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers cites this paper.

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:59:03.324627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T23:07:13.368633Z digest=sha256:5f36b8dfdbbe45d607b9555b913ef19f752c1ac34fd4d0cbbd2a7c79d29c8c3f

Observation 009c2579-6677-4a94-b3b7-aa8f7fa29de0 · inbound

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems cites this paper.

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:45:52.278510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T03:06:29.733729Z digest=sha256:0c401d6439535de218b9e3daf6e05341a227f1b9d4611c99f91786302d6f0378

Observation 46693fd0-7386-4723-be2c-47a2a130340c · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.415052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:494cee338f4afaf64275af002ac179fdedfdcfb5a3a47a6c4c3f8a8fb1fc464d