Pith. sign in

REVIEW 2 cited by

Automatic tagging using deep convolutional neural networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1606.00298 v1 pith:YHA5JIPT submitted 2016-06-01 cs.SD cs.LG

classification cs.SDcs.LG
keywords architecturesautomaticconvolutionaldatasetlayerstaggingarchitecturedifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a content-based automatic music tagging algorithm using fully convolutional neural networks (FCNs). We evaluate different architectures consisting of 2D convolutional layers and subsampling layers only. In the experiments, we measure the AUC-ROC scores of the architectures with different complexities and input types using the MagnaTagATune dataset, where a 4-layer architecture shows state-of-the-art performance with mel-spectrogram input. Furthermore, we evaluated the performances of the architectures with varying the number of layers on a larger dataset (Million Song Dataset), and found that deeper models outperformed the 4-layer architecture. The experiments show that mel-spectrogram is an effective time-frequency representation for automatic tagging and that more complex models benefit from more training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

    cs.SD 2026-07 conditional novelty 7.0 of 10

    State-of-the-art strided audio encoders impose predictable alias-collapse and resolution bottlenecks on frequency primitives; Gabor Latent Refactorization recovers much of the lost separability post-hoc.

  2. Learning Normal Patterns in Musical Loops

    cs.SD 2025-05 reject novelty 4.0 of 10

    A Deep SVDD model using HTS-AT and feature fusion learns normal patterns in variable-length bass and guitar loops, with residual connections improving the learned latent space.

Pith tools