Pith. sign in

REVIEW 2 cited by

Leveraging AI to Generate Audio for User-generated Content in Video Games

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17018 v1 pith:JATP6TIX submitted 2024-04-25 cs.HC cs.AI

classification cs.HCcs.AI
keywords audiocontentuser-generatedgamegamesgeneratorobjectsartificial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In video game design, audio (both environmental background music and object sound effects) play a critical role. Sounds are typically pre-created assets designed for specific locations or objects in a game. However, user-generated content is becoming increasingly popular in modern games (e.g. building custom environments or crafting unique objects). Since the possibilities are virtually limitless, it is impossible for game creators to pre-create audio for user-generated content. We explore the use of generative artificial intelligence to create music and sound effects on-the-fly based on user-generated content. We investigate two avenues for audio generation: 1) text-to-audio: using a text description of user-generated content as input to the audio generator, and 2) image-to-audio: using a rendering of the created environment or object as input to an image-to-text generator, then piping the resulting text description into the audio generator. In this paper we discuss ethical implications of using generative artificial intelligence for user-generated content and highlight two prototype games where audio is generated for user-created environments and objects.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

    cs.SD 2025-07 conditional novelty 6.0 of 10

    AudioBERTScore, a training-free metric combining max-norm and p-norm similarity of audio embeddings, correlates more strongly with human subjective scores for text-to-audio synthesis than conventional metrics.

  2. AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

    cs.SD 2025-05 conditional novelty 5.0 of 10

    AudioTurbo fine-tunes a diffusion model on deterministic noise-audio pairs created by the pretrained Auffusion model, achieving strong text-to-audio results in 10 inference steps.

Pith tools