Pith. sign in

REVIEW 1 cited by

Sound-to-Imagination: An Exploratory Study on Unsupervised Crossmodal Translation Using Diverse Audiovisual Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01266 v2 pith:47T6MDZA submitted 2021-06-02 cs.SD cs.GRcs.MMeess.ASeess.IV

classification cs.SDcs.GRcs.MMeess.ASeess.IV
keywords translationsounddataimagesenoughinformativityperformunknown
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The motivation of our research is to explore the possibilities of automatic sound-to-image (S2I) translation for enabling a human receiver to visually infer the occurrence of sound related events. We expect the computer to 'imagine' the scene from the captured sound, generating original images that picture the sound emitting source. Previous studies on similar topics opted for simplified approaches using data with low content diversity and/or sound class supervision. Differently, we propose to perform unsupervised S2I translation using thousands of distinct and unknown scenes, with slightly pre-cleaned data, just enough to guarantee aural-visual semantic coherence. To that end, we employ conditional generative adversarial networks (GANs) with a deep densely connected generator. Additionally, we present a solution using informativity classifiers to perform quantitative evaluation of the generated images. This enabled us to analyze the influence of network bottleneck variation over the translation, observing a potential trade-off between informativity and pixel space convergence. Despite the complexity of the specified S2I translation task, we were able to generalize the model enough to obtain more than 14%, in average, of interpretable and semantically coherent images translated from unknown sounds.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning

    cs.MM 2024-11 conditional novelty 3.0 of 10

    A review that categorizes deep audio-visual correlation learning methods by architectures, objective functions, datasets, and evaluation metrics, and points to missing standardized benchmarks.

Pith tools