Pith. sign in

REVIEW 2 cited by

Energy Consumption of Deep Generative Audio Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.02621 v2 pith:UKNZHFZN submitted 2021-07-06 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords consumptiondeepenergycommunitygenerativemodelsqualityaudio
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In most scientific domains, the deep learning community has largely focused on the quality of deep generative models, resulting in highly accurate and successful solutions. However, this race for quality comes at a tremendous computational cost, which incurs vast energy consumption and greenhouse gas emissions. At the heart of this problem are the measures that we use as a scientific community to evaluate our work. In this paper, we suggest relying on a multi-objective measure based on Pareto optimality, which takes into account both the quality of the model and its energy consumption. By applying our measure on the current state-of-the-art in generative audio models, we show that it can drastically change the significance of the results. We believe that this type of metric can be widely used by the community to evaluate their work, while putting computational cost -- and in fine energy consumption -- in the spotlight of deep learning research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models

    eess.AS 2025-05 conditional novelty 5.0 of 10

    Inference energy of seven text-to-audio diffusion models grows linearly with denoising steps, while quality saturates, so the best quality-per-energy settings use 10 to 50 steps.

  2. Detecting Musical Deepfakes

    cs.SD 2025-05 conditional novelty 4.0 of 10

    On the FakeMusicCaps benchmark, an ImageNet-pretrained ResNet18 trained on mel spectrograms distinguishes human from AI-generated music with about 88% F1, and stays above 80% F1 on clips with pitch and tempo changes, ...

Pith tools