REVIEW 3 major objections 9 minor 66 references
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
T0 review · 3 major / 9 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Mask-guided image editing can be served far more cheaply by caching the activations of unmasked image regions, cutting average latency by up to 14.7x and raising throughput up to 3x without visible quality loss.
desk verdict Solid serving-systems paper with clean FLOP analysis and real 3x/14.7x wins, but the 'no quality loss' claim rests on thin evidence that needs strengthening. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a block-level activation cache keyed by image template and transformer-block index: for each block, the output activations of unmasked tokens are stored in host memory and loaded into GPU memory ahead of the masked-token computation. Around this cache, the paper wraps four mechanisms: a dynamic-programming pipeline planner that chooses which blocks use cached activations to eliminate loading bubbles; step-level continuous batching that mirrors LLM-style batching but operates on denoising steps and disaggregates CPU-side preprocessing; and linear-regression load models that estimate each worker's latency from mask ratios for greedy request routing. The cache makes each request's FLOPs scale with mask ratio rather than full image size, which is what turns masks into a serving-level optimization.
What would settle it
Compute, for many prompt/mask combinations on the same templates, the cosine similarity between cached unmasked-token activations and activations from a fresh full run; if similarity drops below a threshold in common cases, or if image quality against full regeneration falls below a stated SSIM bar for those cases, the central reuse claim fails there. A concrete protocol: on each supported model, run edits varying prompts and mask shapes, compare the system's output with full regeneration, and check whether SSIM stays above 0.95 throughout.
Extended reading notes
Core claim
The central discovery is that, for a fixed image template, the intermediate activations of unmasked tokens remain highly similar across different edit requests, so they can be cached and reused rather than recomputed. InstGenIE restructures the transformer-block computation: masked tokens are projected into Q, K, V and compute attention among themselves and against the cached unmasked context, while the unmasked tokens' output activations are taken directly from cache. The paper demonstrates the feasibility with a cosine-similarity measurement of activations and attention-score visualizations, then builds the full serving system around it. Evaluated on three diffusion models and production-derived masks, the system reports up to 3x higher engine throughput and up to 14.7x lower average serving latency against the strongest baselines, with SSIM scores of 0.92-0.99 against full regeneration, which the paper reads as visually indistinguishable output.
Load-bearing premise
The load-bearing premise is that activations cached from an earlier request on the same image template remain close enough to what a fresh full-model run would produce for the current prompt and mask, so substituting them does not perceptibly degrade the edited image; this is supported only by a single cosine-similarity measurement on one model and by aggregate FID/SSIM scores, and the paper itself concedes that style-transfer edits would erode the benefit.
Editorial extensions
If this is right
- Editing cost scales with the masked fraction: a mask covering 20% of the image cuts inference work roughly fivefold, so smaller masks yield proportionally larger speedups.
- Template reuse is the precondition: in workloads where the same templates are edited many times, the first edit computes activations that amortize over all later edits.
- Step-level continuous batching stabilizes queueing latency as request rate rises, because new requests join the running batch within one denoising step rather than waiting for the whole batch.
- Mask-size heterogeneity in production traffic requires load balancing at the computation level, not request count or token count, to avoid hot spots.
- The mechanism is model-agnostic across transformer-based diffusion models, so both UNet-based and DiT models can absorb the same caching and scheduling optimizations.
Reading between the lines
- A testable extension would be an adaptive policy that measures cached-activation similarity online and falls back to full computation when similarity drops; the paper does not explore this, but its own single-measurement evidence suggests similarity is the real quality lever.
- The same block-level activation caching could apply to other iterative generative models, such as video diffusion, where static background regions are edited across frames, though cache sizes and attention patterns would need re-checking.
- For workloads with low template reuse, the cache would provide little benefit; the system's reported gains depend on the trace property that roughly 970 templates serve 34 million images, so services with mostly novel inputs would need a different mechanism.
- Global edits like style transfer are the paper's own admitted boundary: because they alter unmasked appearance, the cached activations are no longer valid, so the benefit diminishes exactly where the mask stops being a good specification of the edit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents InstGenIE, a serving system for mask-guided image editing with diffusion models. The core idea is to cache intermediate activations of unmasked tokens from previous requests on the same image template and reuse them, computing only masked tokens. To make this efficient, the paper proposes a bubble-free pipeline that overlaps cache loading with masked-token computation, a continuous batching scheme adapted to diffusion models, and a mask-aware load-balancing scheduler. The evaluation on SD2.1, SDXL, and Flux reports up to 3x higher throughput and up to 14.7x lower average latency versus Diffusers, FISEdit, and TeaCache, with claims of maintained image quality.
Significance. If the central quality-preservation claim holds, the work is significant for production image-editing services: the paper's workload traces show very high template reuse and typically small masks, so the proposed optimizations target a real operational bottleneck. The system design is careful, with a parameter-free FLOP analysis (Table 1), a DP formulation for pipeline scheduling (Algo. 1), and a regression-based scheduler (Algo. 2). The evaluation uses production-shaped workloads and three different model families. The authors also commit to releasing their production trace, which is a valuable community resource. However, the paper's strongest claim—that image quality is maintained—rests on evidence that is currently incomplete, as detailed below. The system-level speedups are plausible but not yet proven to be quality-preserving.
major comments (3)
- [§6.2, Table 2] The claim that InstGenIE 'ensures image quality' or generates images 'visually indistinguishable' from Diffusers requires a direct quality comparison against the Diffusers reference. In Table 2, the Diffusers columns for FID and SSIM are missing for all three benchmarks (SD2.1/InstructPix2Pix, SDXL/VITON-HD, and Flux/PIE-Bench); only CLIP scores are reported for Diffusers on two of the three. Without the Diffusers FID/SSIM values, the InstGenIE numbers (e.g., FID 3.4 for SDXL, SSIM 0.88 for Flux) have no reference point against the standard-quality baseline, and the statement that SSIM 0.99 is 'near-perfect similarity' is not interpretable. The authors should add the missing Diffusers FID/SSIM values, and ideally a Diffusers-vs-Diffusers self-similarity control, to support the equivalence claim.
- [§3.1, Fig. 6] The entire system rests on the assumption that intermediate activations of unmasked tokens, computed for one request on an image template, remain valid for other requests with different prompts and masks. The only direct evidence is the cosine-similarity measurement in Fig. 6, which is reported for SDXL only, with no statement of the number of templates, prompts, mask geometries, or variance across these factors. The caption and text are ambiguous about what is averaged ('average cosine similarities between corresponding masked and unmasked tokens'). Given that attention is global and that prompts can change global appearance, this assumption needs a systematic robustness study with error bars and multiple models. Section 7 concedes that style-transfer edits would erode the benefit, but the paper does not quantify where that boundary lies; this is a load-bearing limitation for the general claim.
- [§4.2, Algo. 1; §4.4, Algo. 2] Algorithm 1 is specified for a single mask ratio m, but the system batches requests with heterogeneous masks and uses the pipeline in that setting. Algorithm 2 refers to a function dp(batch, Comp, Load) that 'extends Algo. 1' to a batch, but the extension is not described. In particular, it is unclear how the cache-loading latency L_m and computation latency C_m are aggregated for a batch containing different masks, and how the block-wise cache decision interacts with requests that join or leave at denoising-step boundaries under continuous batching. Since the throughput and load-balancing claims depend on this batch-level pipeline, the paper should specify the batch-level formulation and validate it separately.
minor comments (9)
- [§2.4] Typo: 'lantecy' should be 'latency'.
- [§1] Typo: 'Adode' should be 'Adobe'.
- [§2.2] Typo: 'editted' should be 'edited'.
- [§6.4] Typo: 'staic' should be 'static'.
- [§6.6] Typo: 'Takeways' should be 'Takeaways'.
- [Table 1] The Cache Shape column is identical for all three operations, which is confusing; the cache is the block output Y, not per-operation caches. Please clarify the notation.
- [Fig. 6] The left panel is labeled 'Similarity' but the y-axis range is 0.5-1.0; the text says 'average cosine similarities between corresponding masked and unmasked tokens' which is ambiguous. Clarify what is being averaged and over how many samples.
- [Fig. 4-Left] The 'Ideal' condition is not defined in the text; please specify what it represents.
- [§6.1] The paper does not specify the exact denoising steps or resolution settings for each model, making the latency results hard to reproduce.
Circularity Check
No material circularity: the central speedup and quality claims are measured against external baselines; only the Table 1 FLOP analysis restates the token-selection rule by construction.
-
self definitional
[Sec. 3.2, Table 1]
"Ori. FLOP O(BLH^2) / Acc. FLOP O(BmLH^2) / Speedup 1/m. In this part, we mathematically analyze the speedup and caching overhead associated with the approach in Fig. 5-Bottom, as summarized in Table 1."
The 'Acc. FLOP' column is defined as the cost after selecting only the masked fraction (mask ratio m) of tokens, and the 'Speedup' column is simply the ratio of the original FLOPs to that reduced FLOPs. Thus the advertised 1/m speedup is a restatement of the algorithm's own token-selection rule rather than an independently derived prediction. The later statement that empirical results are 'consistent with the theoretical analysis in Table 1' (§6.3) is likewise a sanity check of an implementation that by construction skips unmasked tokens. This does not undermine the measured end-to-end latency and throughput comparisons against Diffusers, FISEdit, and TeaCache, but the analytic speedup itself is definitional.
full rationale
The paper's central claims are empirical: end-to-end latency and throughput are measured against external baselines (Diffusers, FISEdit, TeaCache) in Figs. 12 and 14, and image quality is assessed with cosine-similarity measurements (Fig. 6) and FID/SSIM/CLIP comparisons (Table 2). The load-bearing assumption that cached unmasked activations remain valid across requests is supported by a direct cosine-similarity measurement, not by circular reasoning. The regression models used by the mask-aware scheduler are fitted to offline profiled latencies and then used inside the scheduler; this is standard system identification, not a fitted parameter being relabeled as a prediction. Self-citations, notably the Katz trace [37] for the prevalence of image-editing requests, are corroborated by the paper's own independently collected production trace, and no load-bearing uniqueness theorem is imported from prior work by the same authors. The only mildly circular element is the FLOP-count 'analysis' in Table 1, which restates the algorithmic choice to compute only masked tokens; this is a presentational tautology rather than a defect in the experimental validation. Section 7's concession that style-transfer edits would erode the benefit is an honest stated limitation, not evidence of circularity.
Assumptions & free parameters
free parameters (2)
- Compute latency regression coefficients (FLOPs to seconds) =
Not reported; R^2=0.9989 (SDXL), 0.9981 (Flux) on fitted data
- Cache loading latency regression coefficients =
Not reported
assumptions (5)
- domain assumption Unmasked token activations from previous requests are similar enough to fresh activations that reusing them preserves output quality.
- domain assumption The image template used for caching is the same pristine base image across requests; prior edits do not invalidate the cached unmasked activations.
- domain assumption Masked and unmasked tokens attend to each other weakly, so a masked-token-only computation with cached unmasked activations is a good approximation.
- domain assumption Diffusion requests at different denoising steps can be batched together in one model forward pass.
- domain assumption Offline-fitted linear regression models for latency remain accurate for online scheduling decisions.
Cite this review
Pith. "Pith review of InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling." pith.science (2026). https://pith.science/paper/TTAOZADX
@misc{pith2026250520600,
author = {Pith},
title = {Pith review of: InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling},
year = {2026},
howpublished = {\url{https://pith.science/paper/TTAOZADX}},
note = {Machine review of arXiv:2505.20600}
}
read the original abstract
Generative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of masks provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present InstGenIE, a system that efficiently serves image editing requests. The key insight behind InstGenIE is that image editing only modifies the masked regions of image templates while preserving the original content in the unmasked areas. Driven by this insight, InstGenIE judiciously skips redundant computations associated with the unmasked areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, InstGenIE employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, InstGenIE proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogeneous masks induce imbalanced loads, InstGenIE also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, InstGenIE outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3x higher throughput and reducing average request latency by up to 14.7x while ensuring image quality.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
2025. Amazon EC2 P4 Instances. https://aws.amazon.com/ec2/ instance-types/p4/
work page 2025
-
[2]
2025. Amazon EC2 P5 Instances. https://aws.amazon.com/ec2/ instance-types/p5/
work page 2025
-
[3]
Adobe. 2025. Adobe Free Online Photo Editor. https://www.adobe. com/products/photoshop/ai-photo-editor.html
work page 2025
-
[4]
Adobe. 2025. Adobe Free Online Photo Editor. https://www.adobe. com/express/feature/image/editor
work page 2025
-
[5]
Adobe. 2025. Next-level Generative Fill. Now in Photoshop. https: //www.adobe.com/products/photoshop/generative-fill.html
work page 2025
-
[6]
Shubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam, Koyel Mukherjee, and Shiv Kumar Saini. 2024. Approximate Caching for Efficiently Serving Text-to-Image Diffusion Models. In Proc. USENIX NSDI
2024
-
[7]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ram- jee. [n. d.]. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. InProc. OSDI
-
[8]
Friedman, Thomas Williams, Ramesh K
Sohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams, Ramesh K. Sitaraman, and Thomas Woo. 2024. Proteus: A high- throughput inference-serving system with accuracy scaling. InProc. ACM ASPLOS
2024
Show all 66 references
-
[9]
Anyscale. 2025. How continuous batching enables 23x throughput in LLM inference while reducing p50 latency. https://www.anyscale. com/blog/continuous-batching-llm-inference
2025
-
[10]
Stable Diffusion Art. 2023. Adetailer: Automatically fix faces and hands. https://stable-diffusion-art.com/adetailer/
2023
-
[11]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. InProc. NIPS Deep Learning Symposium
2016
-
[12]
Bing-su. 2025. adetailer. https://github.com/Bing-su/adetailer
2025
-
[13]
Tim Brooks, Aleksander Holynski, and Alexei A Efros. [n. d.]. Instruct- pix2pix: Learning to follow image editing instructions. InCVPR
-
[14]
Lequn Chen, Zihao Ye, Yongji Wu, Danyang Zhuo, Luis Ceze, and Arvind Krishnamurthy. 2024. Punica: Multi-tenant LoRA serving. In Proc. MLSys
2024
-
[15]
Seunghwan Choi, Sunghyun Park, Minsoo Lee, and Jaegul Choo. 2021. VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization. InProc. CVPR
2021
-
[16]
Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. 2023. DiffEdit: Diffusion-based semantic image editing with mask guidance. InProc. ICLR
2023
-
[17]
Franklin, Joseph E
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: A low-latency online prediction serving system. InProc. USENIX NSDI
2017
-
[18]
Tri Dao. 2024. FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning. InProc. ICLR
2024
-
[19]
HuggingFace Diffusers. 2025. Create a server. https://github. com/huggingface/diffusers/blob/main/docs/source/en/using- diffusers/create_a_server.md
2025
-
[20]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, and Robin Rom- bach. 2024. Scaling Rectified Flow Transformers for High-Resolution Image...
2024
-
[21]
FastAPI. 2025. FastAPI. https://github.com/fastapi/fastapi
2025
-
[22]
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024. Cost- Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention. InProc. ATC
2024
-
[23]
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kauf- mann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving DNNs like Clockwork: Performance predictability from the bottom up. InProc. USENIX OSDI
2020
-
[24]
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thi- nakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R. Das
-
[25]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. InProc. EMNLP, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (Eds.)
2021
-
[26]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. InProc. NIPS
2017
-
[27]
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InProc. ICLR
2022
-
[28]
HuggingFace. 2025. Accelerate inference of text-to-image diffu- sion models. https://huggingface.co/docs/diffusers/en/tutorials/fast_ diffusion
2025
-
[29]
HuggingFace. 2025. Inpainting. https://huggingface.co/docs/diffusers/ en/using-diffusers/inpaint
2025
-
[30]
Asynchronous I/O. 2025. Asynchronous I/O. https://docs.python.org/ 3/library/asyncio.html
2025
-
[31]
Xuan Ju, Ailing Zeng, Yuxuan Bian, Shaoteng Liu, and Qiang Xu. 2024. PnP Inversion: Boosting Diffusion-based Editing with 3 Lines of Code. InICLR
2024
-
[32]
Jeongho Kim, Guojung Gu, Minho Park, Sunghyun Park, and Jaegul Choo. 2024. Stableviton: Learning semantic correspondence with latent diffusion model for virtual try-on. InProc. CVPR
2024
-
[33]
Conference’17, July 2017, Washington, DC, USA
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Conference’17, July 2017, Washington, DC, USA
2017
-
[34]
Black Forest Labs. 2024. FLUX. https://github.com/black-forest-labs/ flux
2024
-
[35]
Muyang Li, Tianle Cai, Jiaxin Cao, Qinsheng Zhang, Han Cai, Junjie Bai, Yangqing Jia, Ming-Yu Liu, Kai Li, and Song Han. 2024. DistriFusion: Distributed parallel inference for high-resolution diffusion models. In Proc. IEEE/CVF CVPR
2024
-
[36]
Suyi Li, Hanfeng Lu, Tianyuan Wu, Minchen Yu, Qizhen Weng, Xusheng Chen, Yizhou Shan, Binhang Yuan, and Wei Wang. 2025. Toppings: CPU-Assisted, Rank-Aware Adapter Serving for LLM Infer- ence. InProc. USENIX ATC
2025
-
[37]
Suyi Li, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Zhipeng Di, Weiyi Lu, Jiawei Chen, Kan Liu, Yinghao Yu, Tao Lan, Guodong Yang, Lin Qu, Liping Zhang, and Wei Wang. 2025. Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters. InProc. USENIX ATC
2025
-
[38]
Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, and Lili Qiu. 2024. Parrot: Efficient Serving of LLM-based Applications with Semantic Variable. InProc. OSDI
2024
-
[39]
Feng Liu, Shiwei Zhang, Xiaofeng Wang, Yujie Wei, Haonan Qiu, Yuzhong Zhao, Yingya Zhang, Qixiang Ye, and Fang Wan. 2025. Timestep Embedding Tells: It’s Time to Cache for Video Diffusion Model. InProc. IEEE/CVF CVPR
2025
-
[40]
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun- Yan Zhu, and Stefano Ermon. 2022. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations. InProc. ICLR
2022
-
[41]
Midjourney. 2025. Editor - Midjourney. https://docs.midjourney.com/ hc/en-us/articles/32764383466893-Editor
2025
-
[42]
OpenAI. 2025. OpenAI CLIP. https://huggingface.co/openai/clip-vit- base-patch16
2025
-
[43]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2024. SDXL: Improving Latent Diffusion Models for High-Resolution Image Syn- thesis. InProc. ICLR
2024
-
[44]
Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2025. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot. InProc. FAST
2025
-
[45]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proc. ICML
2021
-
[46]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. InProc. IEEE/CVF CVPR
2022
-
[47]
Gonzalez, and Ion Stoica
Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, Joseph E. Gonzalez, and Ion Stoica. 2023. S-LoRA: Serving thousands of concurrent LoRA adapters. InProc. MLSys
2023
-
[48]
Gonzalez, and Ion Stoica
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. 2024. Fairness in Serving Large Language Models. InProc. OSDI
2024
-
[49]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. At- tention is all you need. InProc. NIPS
2017
-
[50]
Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, Dhruv Nair, Sayak Paul, William Berman, Yiyi Xu, Steven Liu, and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/ diffusers
2022
-
[51]
Luping Wang, Lingyun Yang, Yinghao Yu, Wei Wang, Bo Li, Xianchao Sun, Jian He, and Liping Zhang. 2021. Morphling: Fast, near-optimal auto-configuration for cloud-native model serving. InProc. ACM SoCC
2021
-
[52]
Qixun Wang, Xu Bai, Haofan Wang, Zekui Qin, Anthony Chen, Huaxia Li, Xu Tang, and Yao Hu. 2024. InstantID: Zero-shot identity- preserving generation in seconds.arXiv preprint arXiv:2401.07519 (2024)
2024 arXiv
-
[53]
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023. Tabi: An efficient multi-level inference system for large language models. In Proc. ACM EuroSys
2023
-
[54]
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli
-
[55]
Bingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024. dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving. InProc. USENIX OSDI
2024
-
[56]
Yuhao Xu, Tao Gu, Weifeng Chen, and Arlene Chen. 2025. OOTDiffu- sion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-On.Proc. AAAI(2025)
2025
-
[57]
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A distributed serving system for transformer-based generative models. InProc. USENIX OSDI
2022
-
[58]
Zihao Yu, Haoyang Li, Fangcheng Fu, Xupeng Miao, and Bin Cui. 2024. Accelerating text-to-image editing via cache-enabled sparse diffusion inference. InProc. of AAAI
2024
-
[59]
ZeroMQ. 2025. ZeroMQ. https://github.com/zeromq/pyzmq
2025
-
[60]
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019. MArk: Exploiting cloud services for cost-effective, SLO-aware machine learn- ing inference serving. InProc. USENIX ATC
2019
-
[61]
Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica. 2023. Shepherd: Serving DNNs in the wild. InProc. USENIX NSDI
2023
-
[62]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding Con- ditional Control to Text-to-Image Diffusion Models. InProc. IEEE/CVF ICCV
2023
-
[63]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. InProc. NIPS
2024
-
[2004]
Image Process.(2004)
Image quality assessment: From error visibility to structural similarity.IEEE Trans. Image Process.(2004)
2004
-
[2022]
Cocktail: A multidimensional optimization for model serving in cloud. InProc. USENIX NSDI
-
[2023]
Efficient Memory Management for Large Language Model Serv- ing with PagedAttention. InProc. SOSP
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.