Pith. sign in

REVIEW 1 cited by

Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11267 v2 pith:G5IPWDGB submitted 2023-04-21 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords modelsdiffusionlargedeviceson-deviceoptimizationsuserability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development and application of foundation models have revolutionized the field of artificial intelligence. Large diffusion models have gained significant attention for their ability to generate photorealistic images and support various tasks. On-device deployment of these models provides benefits such as lower server costs, offline functionality, and improved user privacy. However, common large diffusion models have over 1 billion parameters and pose challenges due to restricted computational and memory resources on devices. We present a series of implementation optimizations for large diffusion models that achieve the fastest reported inference latency to-date (under 12 seconds for Stable Diffusion 1.4 without int8 quantization on Samsung S23 Ultra for a 512x512 image with 20 iterations) on GPU-equipped mobile devices. These enhancements broaden the applicability of generative AI and improve the overall user experience across a wide range of devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions

    cs.SD 2025-01 conditional novelty 6.0 of 10

    A single audio-plus-text LLM jointly performs voice trigger detection, device-directed speech detection, dialog act classification, and ASR, with reported EER reductions of 64% and 22% over dedicated baselines.

Pith tools