Pith. sign in

Paper Citation Record · LEDGER

DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2207.00032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2207.00032 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:37:14.126135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T22:59:03.322421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39900118-976a-4971-8a9c-d9c4cbe878d9 · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.431176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:fab18b76c908b25e13eaa998b502769bb76be2ccc9e97a2f4b99518bbc0d93c3

Observation 08dccdd5-98d0-49e9-8fc2-68424e18710b · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.819111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:acd02d7d944c9e52e1c367482c11fbe90ef4966b49855fefa87c42352d5b9d99

Observation d873e317-48e0-4e95-9548-57e999c07e5a · inbound

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference cites this paper.

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:01:46.202032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:01:46.202032Z digest=sha256:8c1346c2f8d35ce24c9c03cae62844edfe947e7a54c5023f4bd95494cbcf5054

Observation 77777fa4-ed36-491d-ae1c-725ae2020a07 · inbound

iServe: An Intent-based Serving System for LLMs cites this paper.

iServe: An Intent-based Serving System for LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:14.126135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:14.126135Z digest=sha256:d6e81d4ee54d7131ced1bf0b6a04ececad4a534f3ffe1156438e7062ae8442f5

Observation dd5d2662-76b1-44e0-91a3-18f9d06fb1ef · inbound

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making cites this paper.

Aero-LLM: A Distributed Framework for Secure UAV Communication and Intelligent Decision-Making DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:18:07.705592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:18:07.705592Z digest=sha256:30939770e8a7b59ca210209522ec006d7cc4b745a56fe6870f7e4d842d19a04a

Observation 58157a7a-196f-4952-a0fe-cc8c9c97bf12 · inbound

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation cites this paper.

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:16.021201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:16.021201Z digest=sha256:de2e282b20750a4926f056e3fb0f91e71a47a99295f73d3e0c23714affa8155b

Observation c8abc7c3-bda3-4139-8f81-e0cb39023140 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:29.198408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:29.198408Z digest=sha256:d10a7710546117f77e0f530a6aa3ff8d88ebb8b95355a5504b3153a0a42ecf5d

Observation 2f5b486c-3f15-42ee-8066-5e3bcf0cabab · inbound

EvolveSearch: An Iterative Self-Evolving Search Agent cites this paper.

EvolveSearch: An Iterative Self-Evolving Search Agent DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:55.546676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:55.546676Z digest=sha256:e12e17178b930be32920a59fb6b6bf1924b89168a6381f3b326563adb14b26ca

Observation 1bc7a03b-05fc-4f7c-abe1-632d5628a92a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.268813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:cf1b408d9cb75d0bc96ff7c0069e3f6b2b59b9daa4b4f82da4939cbd9f9840ad

Observation de56e279-cca7-41aa-b6f5-8e0a892fad08 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:29.589719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:a507cafe52e356c214b682719b57cbd56bfb3d6fd0f1c24bfef1cd907cb1a0dd

Observation 48682b07-9fa5-4f44-9259-f8e962a6d74b · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.439048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:448c2cf7cd705dceecb57a66f66abc644ac7e140f05a3bd5dbd0045b34c0d83e

Observation 97c1c357-ece0-4476-ad31-83fd25f2782d · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:41:08.897769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:18:05.789417Z digest=sha256:b858b9d88ea390b04848f418e512eead2f4c802265f97d9e92233cdc36f8215b

Observation 4b7cea5c-8342-4858-b4d1-0379d2ba33f4 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:49:48.919062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T06:49:16.271699Z digest=sha256:7398aca908a2f627b9947018fd01b7b9e2ebe295ca865adb35f7e05796a942c2

Observation e122ca3f-c6ac-4e43-8450-0c2acab501e7 · inbound

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs cites this paper.

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:23:43.898162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:23:43.898162Z digest=sha256:a78f02b231a938744694a74e3efa4724246c5611d3c45057e0e7dc83b8b614cd

Observation bafbf0de-24ce-459a-8c15-d717f12901b7 · inbound

ShardTensor: Domain Parallelism for Scientific Machine Learning cites this paper.

ShardTensor: Domain Parallelism for Scientific Machine Learning DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:07:09.104038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T03:03:32.316700Z digest=sha256:6aea02a333a26530a20d350141744e298a4b4ec92b80fef1d3a70ada77e6397d

Observation a83bf6aa-9455-413c-b910-edc84e1334a9 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.365091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:457f22bdcaf7e16a222ca27387a74e7c6386470efb4e7366818c8c76eb255c57

Observation 0bf0fca9-42f7-4096-af22-a01062156484 · inbound

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving cites this paper.

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:03:47.843565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T18:03:19.798168Z digest=sha256:0cd247e351090e28d9acf778c8ee141386de5025fe6f97a737f1ee2a4244f1ee

Observation 2ca665c7-8363-4c6b-8dd6-93f5aa9b6272 · inbound

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers cites this paper.

AoiZora: Topology-Aware Auto-Parallel Optimization for Inference of Diffusion Transformers DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:59:03.324627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T23:07:13.368633Z digest=sha256:77828e690979d70a2e9bda922bb9f4ec995e04347ae5456d7fcff2813127165c

Observation 009c2579-6677-4a94-b3b7-aa8f7fa29de0 · inbound

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems cites this paper.

Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:45:52.278510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T03:06:29.733729Z digest=sha256:85a9d2b9cbd6a08957dd1e0d07b4133fd229813a9b404cab6739c07950d55a97

Observation 46693fd0-7386-4723-be2c-47a2a130340c · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.415052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:699d44f3fdf3cf46421392851457d9d771e4636e4fd992a3137b397ca64daa3c