Pith. sign in

Paper Citation Record · LEDGER

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 4 inbound Pith citation observations for arXiv:2411.15381.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15381 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:27:01.856685Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:49:04.349040Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact2
  • verified fuzzy28
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6b339838-4b1f-4ca5-8aef-35986f752893 · outbound

This paper cites https://archive.org/details/archiveteam-twitter-stream-2018-04, 2018.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling https://archive.org/details/archiveteam-twitter-stream-2018-04, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.893254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.595608Z digest=sha256:62e95d6c7714931d217f798857dbef0ae9e383271b8af1c8fc1ce746f4cb5407

Observation 3f516ddc-f2b9-4a5d-bab4-6bcc75ed5b41 · outbound

This paper cites build, train, and deploy machine learning models at scale.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling build, train, and deploy machine learning models at scale

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.876943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.601213Z digest=sha256:9fb2f0b0f3df666ea08829e309391146051d0770b4c1d5783a048dc7cf9c6ee4

Observation 8f690e63-0b8b-462d-bf61-ea9fd31dfbec · outbound

This paper cites https://developer.nvidia.com/nvidia-triton-inference-server, 2022.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling https://developer.nvidia.com/nvidia-triton-inference-server, 2022

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.859926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.606258Z digest=sha256:8baeaa2a05c21faae645ffc5f4455795e573c638d0d8ccbde314e82530ea901f

Observation 86fedf91-a8dc-47e3-aea6-4273b14aee71 · outbound

This paper cites Adobe firefly.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Adobe firefly

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.843476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.611958Z digest=sha256:2abcef88b7fbb63b45e12ee1f995ab9a216503dd09fa3e57ebdb94692ad5ec0a

Observation 58999957-9aff-4a58-880e-3deefce50e6a · outbound

This paper cites D., Williams, T., Sitaraman, R.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling D., Williams, T., Sitaraman, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.826968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.617052Z digest=sha256:ce8f33aa9b6108f36916a88758843875fcefd115697a4671027520ea0384ecdd

Observation ec542047-654e-44bc-b8c1-dbe7c2be5ffe · outbound

This paper cites Batch: Machine learning inference serving on serverless platforms with adaptive batching.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Batch: Machine learning inference serving on serverless platforms with adaptive batching

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.810385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.622178Z digest=sha256:1923a454304d58b5acde49c678ed79c6182d7c2fe510edde57fc4bed978a1eac

Observation c8ff1979-f004-4dfa-bb97-b99e5955f732 · outbound

This paper cites Adaptive neural networks for efficient inference.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Adaptive neural networks for efficient inference

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.792984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.627824Z digest=sha256:b774c066624d435e352d2b4583a645d6ed94f9dc86503544716c0bcacb6143b6

Observation 93b30f8b-798e-404c-b557-c20d29f92779 · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.632654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.632654Z digest=sha256:915f10242ee07a003795bf3944095ac9c3c656555a0a6167a9bdaf1f44dbfd6b

Observation d2c6f702-a993-49e4-8ad1-dd4d94f6d50c · outbound

This paper cites J., Gonzalez, J.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling J., Gonzalez, J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.775787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.637742Z digest=sha256:e7d08eabae5372d339d86a3cb35d6849fdcfadae2e6e645c8351a7eb5116594f

Observation 1af2c845-98d8-4cb8-bd55-9faa881b24c6 · outbound

This paper cites J., Gonzalez, J.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling J., Gonzalez, J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.758588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.642574Z digest=sha256:f3b61350c1f4c96637b2ddf1a4acdbf358b8da669bcece11617d455b60998e0d

Observation 5511b38f-1a35-4125-8dab-0aeb9a02a198 · outbound

This paper cites Inferline: latency-aware provisioning and scaling for prediction serving pipelines.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Inferline: latency-aware provisioning and scaling for prediction serving pipelines

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.742137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.647120Z digest=sha256:0136a3b0d70978c76cff0e2e980b060b87d279d649d4664849475f41e2487757

Observation e5d89cde-ccfc-4b19-bac1-99bf2d7ce8c4 · outbound

This paper cites Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.652242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.652242Z digest=sha256:a7f82c85aabfcebe0655773f1eaafc3e138b18ebd3877ee791eccc16d42565bc

Observation 2b3756cf-e1f7-4930-b950-1b63b1470ed1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.657683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.657683Z digest=sha256:51738cc71ef6bc8200afec9ea8157ddcc4fce02e0b400ab961c618c21a55bfed

Observation 8d171270-cf19-41ae-9fa3-1077eb2cbab0 · outbound

This paper cites Serving \ DNNs \ like clockwork: Performance predictability from the bottom up.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Serving \ DNNs \ like clockwork: Performance predictability from the bottom up

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.726100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.663000Z digest=sha256:a5a642bfa1341c18edfde58ccd2f3626a99aebbaba4532824c1650cf86fd2a43

Observation c6e34ff0-4642-4fa4-a7b2-cb0676c77479 · outbound

This paper cites R., Mishra, C.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling R., Mishra, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.708691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.667892Z digest=sha256:6bc66c38a307a9424ded77bedd70b80d9d081216fc7abbdb44e0e5e4b890335a

Observation 8bead135-69f7-4fc0-bca4-b0912a73c18e · outbound

This paper cites Sommelier: Curating dnn models for the masses.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Sommelier: Curating dnn models for the masses

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.690982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.672853Z digest=sha256:ecf9ef060df67030f771fbe59f7b9f7e6b67889121b941c6190303cee59df631

Observation afa21dde-a470-43fe-bd5e-44e1adb90d51 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Language Model Cascades: Token-level uncertainty and beyond

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.678122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.678122Z digest=sha256:36eb0ad84bedfa9e52a6765069d4eafa1c52452615cecf1cfb2d6b16e27c2f82

Observation fb15b216-d4a8-4b73-8447-d9d70d83ae09 · outbound

This paper cites URL https://www.gurobi.com/.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling URL https://www.gurobi.com/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.674398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.683265Z digest=sha256:cd277fe21b183d6cf8a1d9c4c9150a79b60482cd5746b7e0d881174034e02c35

Observation 831dba85-a24b-47af-90bc-7754b85066ce · outbound

This paper cites Deep Residual Learning for Image Recognition.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Deep Residual Learning for Image Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.688497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.688497Z digest=sha256:391474e516e9854c46e21f0ceb8ff781cd742ec6e9106a6e59cee1d8cd6d2920

Observation 596d9d61-41ed-4515-be45-4f5c66affc04 · outbound

This paper cites CLIPS core: A reference-free evaluation metric for image captioning.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling CLIPS core: A reference-free evaluation metric for image captioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.693864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.693864Z digest=sha256:d47a59eab4668042dab865e207b32144184856a39811f3cfa575b93acaec545e

Observation 221f98b3-6b47-4258-b385-46660cb19b42 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.657412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.699078Z digest=sha256:c87f5f9f84aaf21e1c234afcdc5eafdeb2769279d0d7d45f6931c2f4b1820bd9

Observation 675856fc-7aa9-4be7-a41c-f922385c0ffc · outbound

This paper cites Scrooge: A cost-effective deep learning inference system.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Scrooge: A cost-effective deep learning inference system

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.640758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.703949Z digest=sha256:8ff6ebc38757eb06347a56d5df58b4798b32c11dad62e040fd5fa12ef854e922

Observation 8e738f5b-25ea-4398-8501-25fd38922ff3 · outbound

This paper cites Hugging face models hub, 2024.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Hugging face models hub, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.624099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.708704Z digest=sha256:752f9ddef99a2908e22147c9c6ecd8613d7116de3c76707b508f2f7f12993e7b

Observation ccba5706-5591-4be1-ad01-5bbfc39aed41 · outbound

This paper cites Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.713566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.713566Z digest=sha256:5a58d9a04bbbe4a1cd5cafdc8878306359981aa1a4015422c0251b1022315749

Observation 6073e157-e86f-4063-9833-cfbb4dcdb6c7 · outbound

This paper cites Willump: A Statistically-Aware End-to-end Optimizer for Machine Learning Inference.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Willump: A Statistically-Aware End-to-end Optimizer for Machine Learning Inference

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:27:02.224872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.718943Z digest=sha256:84c7a4eaae66349d5360a4f2004be7819586a2a2543e5ba03b871248d3a456c3

Observation e6bee921-81a2-4d50-9fd9-b0c18ccd2fab · outbound

This paper cites CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.724169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.724169Z digest=sha256:717a6e9cc3ef11762316aa7134dcec89deaff0307852dd313156aeeb39b3f60e

Observation ebc4d86d-2161-4947-a409-a2d6f5d06e81 · outbound

This paper cites SDXL-Lightning: Progressive Adversarial Diffusion Distillation.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.729111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.729111Z digest=sha256:d9804b1e0e43485498a2967ffbb3fef3e8cf407f26d0f5987e42a500307b2043

Observation 071f4e84-85a7-4aa1-9acc-20a0a828fe2a · outbound

This paper cites an unresolved cited work.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.734396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.734396Z digest=sha256:1b1cb4360256f215af59777dc6dc13a44614d3d3f594ec99fade200de509a541

Observation 1a7ed603-b3e1-4db8-8b5e-54a276106107 · outbound

This paper cites Midjourney.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Midjourney

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.595379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.739242Z digest=sha256:57fbfe93670b2ff3a2f8248c9f1a1f401f42b3f80849985fb937186c8bbf9a6e

Observation 61a01d18-ab89-45a8-b837-01d7674aabc9 · outbound

This paper cites Tensorflow-serving: Flexible, high-performance ml serving.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Tensorflow-serving: Flexible, high-performance ml serving

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.576033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.743995Z digest=sha256:305380204af240aab45c203608ccc3b560441e665d65a23fe02162480cb5d5ce

Observation 408dae85-b0fd-44d7-ae8f-c8947e41f859 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.749149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.749149Z digest=sha256:85f340e5931e33fdce6a32f6bfa3ab43011a50ea9d51abfcc6476070cc65d91c

Observation 3743c3a7-a30f-4d40-a3b8-f1801f5fa9f3 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.754594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.754594Z digest=sha256:7af988d485c920f34a334cd22351c42841da477e0e758a7abfd27faa92ce16af

Observation d3e8aab3-6934-4a2f-a9cd-e867cf47761f · outbound

This paper cites Torchserve.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Torchserve

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.559073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.760153Z digest=sha256:66c1417d4b266721d4ccdc5dfab8ee7a08d87bfd23ec64cb134ae5fe80d61a98

Observation 1f1eb2e6-3714-460f-9e35-e3e656c2b8d2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.764989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.764989Z digest=sha256:a28773ebd221af51da24526c4d758011ce05b370b6f913fc5de5f3097b61047c

Observation e847266e-ed6b-4b24-96ac-0a05db6a3a9e · outbound

This paper cites J., and Kozyrakis, C.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling J., and Kozyrakis, C

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.530876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.769792Z digest=sha256:8f40613acd82f2b1cb302fe3ee0201b29682296bf28b576c3402e1a9806953e6

Observation 9387b0e1-0e28-4c43-9032-838476b8398c · outbound

This paper cites J., and Kozyrakis, C.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling J., and Kozyrakis, C

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.774915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.774915Z digest=sha256:4888f9f6c3a8db7c6481f759c8f0ece8b5daae0370ee691ef177f6f85ee28f77

Observation 5b040390-ec84-4503-bc81-f202af14a8af · outbound

This paper cites Adversarial Diffusion Distillation.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Adversarial Diffusion Distillation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.779693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.779693Z digest=sha256:6ec6238901943f98ab863a3e6ffc6de051af16ae301af92f48a26d23a2392ea9

Observation c59f4356-5ad6-4298-b962-a5860d120513 · outbound

This paper cites Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.512873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.784847Z digest=sha256:65c23636a9f034ac8f8433ab5fbec01a21a72c07176585fa5008d5e826aeb140

Observation d815e354-5d72-437f-8a28-1efbf8d16643 · outbound

This paper cites Fast video classification via adaptive cascading of deep models.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Fast video classification via adaptive cascading of deep models

Reference 39

Resolution
verified exact
doi, observed 2026-08-12T14:27:01.899001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.789531Z digest=sha256:205dab7fb29e59b4aa7750e5981819775f43f3852001c933a44c925331a43061

Observation 19ee7a8c-2775-49ec-a92d-3270d01c7919 · outbound

This paper cites Nexus: A gpu cluster engine for accelerating dnn-based video analysis.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Nexus: A gpu cluster engine for accelerating dnn-based video analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.495517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.794294Z digest=sha256:bbe186bf7dd8cad70f51af71918b33990be5c2a595c2072eade10fa6d1d55928

Observation ff5a7a63-9d26-40fb-b520-776126789cc3 · outbound

This paper cites F., Thompson, J.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling F., Thompson, J

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.476846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.799021Z digest=sha256:3f868cc0d2e65f83d261ec7d605631f0bcebe6a63e7867133c46ea485d6d03da

Observation c13e7ee5-d1f4-47af-ad5f-386596fee70a · outbound

This paper cites SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling SDXS: Real-Time One-Step Latent Diffusion Models with Image Conditions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.804062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.804062Z digest=sha256:306170affa98da76c30778d994b119a31b4ff7dac9b0dc83e53a7f3003e79e5c

Observation 13a5051f-25bc-4043-a85b-7fef7d6c38b6 · outbound

This paper cites Tensorflow serving.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Tensorflow serving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.458559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.809350Z digest=sha256:ca6723b87bfd8c40a4407892e7b9cb98f66617ef47d6244abe8fc77cd166717e

Observation 3c9d8c47-a515-4899-88e0-8d96b6993664 · outbound

This paper cites and Jones, M.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling and Jones, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.439690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.814699Z digest=sha256:0f4b1a9a7ffc09efff1a6e4173ed2a1e15acb21818a9065826c43d079dbdc109

Observation 59fbc0a5-2a44-49ed-819a-3e18267f3e40 · outbound

This paper cites Diffusers: State-of-the-art diffusion models.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Diffusers: State-of-the-art diffusion models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.819689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.819689Z digest=sha256:25e3c8172b3057e29b3d8633149bc8140805d115a8c5b78b82510487dc3a61e0

Observation adb7d2c2-4ffe-401a-b5ef-2a1263f5640a · outbound

This paper cites Rafiki: Machine Learning as an Analytics Service System.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Rafiki: Machine Learning as an Analytics Service System

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.824439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.824439Z digest=sha256:839127d5bfadb7dc3257b5acd72579172c432d0fb2766a44e19af4364dbe4fe3

Observation 89663e8d-cda5-4c9d-9623-e969ec33819b · outbound

This paper cites Tabi: An efficient multi-level inference system for large language models.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Tabi: An efficient multi-level inference system for large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.830067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.830067Z digest=sha256:63f241357b1748f6bb63ec5e37962356ecfc36e363050d2fc4805673bc916f1b

Observation 95183afe-e1c2-4a92-8924-015001f28317 · outbound

This paper cites DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.835347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.835347Z digest=sha256:649fe4b146ccccddcfd4292029e2ca749b5c197838182fca3fc3c9dcf99e800b

Observation ef582d90-0749-4755-ae06-615f15cc131c · outbound

This paper cites Bartscore: Evaluating generated text as text generation.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Bartscore: Evaluating generated text as text generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.410087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.841104Z digest=sha256:5c65d61a3f7d960fa2dab0cbf667e8cf9f37136e70edf5f555c53ae8351d2e41

Observation 009eb8c5-4744-4047-8ab1-5ddf6b08bfd3 · outbound

This paper cites an unresolved cited work.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:27:02.392540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.846613Z digest=sha256:ad5065a143a08e24e497824fb06e03d91f7d0dead933c7e884a667de39ac908f

Observation 9e38a922-a2a5-46c9-bd5c-f1ce786729f0 · outbound

This paper cites \ Model-Switching \ : Dealing with fluctuating workloads in \ Machine-Learning-as-a-Service \ systems.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling \ Model-Switching \ : Dealing with fluctuating workloads in \ Machine-Learning-as-a-Service \ systems

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:27:02.374533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T14:27:01.851437Z digest=sha256:63bb5ce9a82db93c60dd62851a10e470318efb89070ef570bcf608c34094d126

Observation 7de3969a-5eb3-4168-9beb-f383f9ca6902 · outbound

This paper cites write newline.

DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:27:01.856685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:27:01.856685Z digest=sha256:ca2a8852020bd063e992b1d7778bda49b8b165827c61e4511dc4eb4ad8dd4d6f

Pith citing papers

Observation f91bb7ea-c123-4962-a7a1-6a9b52d1bebf · inbound

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving cites this paper.

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:53:57.613373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T20:46:51.896596Z digest=sha256:4c8666b789c1959c937b37e589a5b8dba3b06c03484df306196b3cbf0310e308

Observation 245aba40-2b5f-4cbe-849f-fc9060b86073 · inbound

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving cites this paper.

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.323307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T05:35:01.339479Z digest=sha256:1960e587557eb66d320213f907d29b8d15edf981e102eda929f5be8ccac90586

Observation 40dac2ac-8156-42d6-8f8f-9a1b85da25d0 · inbound

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving cites this paper.

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.300197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T23:50:39.241880Z digest=sha256:065b9ca46cbed206c674f96e4ec01f5e1f949162b946b31ddc8a50e2d24c5749

Observation 9f45b247-6672-4d76-81e7-15703380acc9 · inbound

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration cites this paper.

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T06:49:04.349040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:49:04.349040Z digest=sha256:c930e193ffba133f9031ce27c9c78922058063feb33b4c28e1252fcb31bc1cfd