REVIEW 9 cited by
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive Architecture (A-JEPA), a simple extension method for self-supervised learning from the audio spectrum. Following the design of I-JEPA, our A-JEPA encodes visible audio spectrogram patches with a curriculum masking strategy via context encoder, and predicts the representations of regions sampled at well-designed locations. The target representations of those regions are extracted by the exponential moving average of context encoder, \emph{i.e.}, target encoder, on the whole spectrogram. We find it beneficial to transfer random block masking into time-frequency aware masking in a curriculum manner, considering the complexity of highly correlated in local time and frequency in audio spectrograms. To enhance contextual semantic understanding and robustness, we fine-tune the encoder with a regularized masking on target datasets, instead of input dropping or zero. Empirically, when built with Vision Transformers structure, we find A-JEPA to be highly scalable and sets new state-of-the-art performance on multiple audio and speech classification tasks, outperforming other recent models that use externally supervised pre-training.
Forward citations
Cited by 9 Pith papers
-
Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture
A JEPA encoder finetuned on synthetic bar charts enables a lightweight decoder to recover bar values from chart images, but the method remains behind state-of-the-art supervised systems.
-
The JEPA Paradox in Language: The Geometry of Linguistic Alternatives
Pure squared-error JEPA on masked text saturates mutual information and keeps high conditional variance, then collapses in rank and cosine similarity and transfers poorly, unlike matched I-JEPA on images.
-
Music-JEPA: Learning a World Model of Sound from Action
An action-conditioned JEPA trained on paired piano audio and MIDI learns latent sound dynamics that support MIR tasks and transcription-style planning.
-
Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
JEPA-based self-supervised pre-training on inertial sensor data improves recognition of rare transitional human activities compared to supervised learning, with gains mostly on transition classes.
-
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
Predicting the outputs of several hidden layers of an EMA teacher, rather than only the final layer or pixels, substantially improves self-supervised ViT representations on ImageNet and downstream tasks.
-
Discrete JEPA: Learning Discrete Token Representations without Reconstruction
Discrete-JEPA learns discrete semantic image tokens through latent predictive coding without pixel reconstruction, and achieves stable long-horizon prediction on synthetic symbolic tasks.
-
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
SSLAM pre-trains audio transformers on partially mixed audio clips with a source retention loss, improving polyphonic sound tagging while keeping monophonic benchmark scores.
-
Joint-Embedding Predictive Architecture for Solar PV Panel Fault Classification
JEFFNet fuses StoP-JEPA semantic embeddings with EfficientNetV2-S features for thermal IR PV fault classification, beating GEPFNet on F1 for multiclass and binary tasks with 47% fewer parameters.
-
Scalable and Efficient Joint Spiking Embedding Predictive Architecture for Large-Scale Dynamic Graphs
SG-JEPA applies joint-embedding predictive learning to dynamic graphs, using spiking-neuron context encoders to predict future node embeddings without edge reconstruction or graph augmentation.
Discussion (0). Sign in to comment.