Pith. sign in

Paper Citation Record · LEDGER

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

As of 23 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 9 inbound Pith citation observations for arXiv:2507.02915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02915 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:56:46.342431Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:15:06.622581Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:12.629815Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact12
  • verified fuzzy8
  • unresolved20
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f1a916e-e772-4f3d-8282-a493595b354f · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.224028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.224028Z digest=sha256:fecedd01f6d631369ba273e88ee40a215ddb2790c6e295648f347415a51c33bb

Observation 9ceb76b9-7dbc-4c05-9baa-86471b24b068 · outbound

This paper cites CED: Consistent ensemble distillation for audio tagging.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CED: Consistent ensemble distillation for audio tagging

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.266196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.227882Z digest=sha256:c0d245717bd20b004f48465aa014952c270af2b182db24794fbe85c965282958

Observation b46a13a5-5276-415c-a05e-6a163b4ef094 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Scaling up masked audio encoder learning for general audio classification

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.230938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.230938Z digest=sha256:e6b7d9db49c1b2636f082ddefa85f92a91e3efb6e19129e1eda31cbc96977c37

Observation 79718107-558b-4fc2-916e-de189db3939b · outbound

This paper cites Masked Modeling Duo: Towards a Universal Audio Pre-training Framework.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Modeling Duo: Towards a Universal Audio Pre-training Framework

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.251153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.234160Z digest=sha256:86894febbd3bd0b3f7c3ee914625a7010a872d3f7a7de19cc096902c96424340

Observation fb09b00d-30d2-47af-9ac9-1eb9bb433875 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.237565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.237565Z digest=sha256:dd3cf77feecf9faa3e65b7173b7e6964b7a530ec860026611b04b4bf7e6a1941

Observation 8461dad3-f481-4f9a-8785-b4b94006be8c · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.241459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.241459Z digest=sha256:41070530e984192c32b208e78b57190d519d9208957aad85c4ce2e92c2acb51a

Observation 4615b1ab-4aa0-48a8-b010-a13417829438 · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.244140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.244140Z digest=sha256:199459f6a27dbdd444f46797ff041869d6e6e02d04a22e3f1cfa69af30ce2fe1

Observation 148ab8b8-5667-4689-ba24-ac83137fca14 · outbound

This paper cites Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Efficient Self-super- vised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.339338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.247277Z digest=sha256:632d3677af10e17415615e34af52cd880d472b2d485804c55742943a689c38e8

Observation e1896868-56f7-43f6-85f3-e72bd47c4294 · outbound

This paper cites A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.332558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.249998Z digest=sha256:a3f1735c1c2cec7ea019bcc0eee6353b2ce7351e1f20bf0853b4af53dca853ba

Observation 02ad345b-f45c-48aa-a64c-c36c78f35ebb · outbound

This paper cites Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.252958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.252958Z digest=sha256:e8695f9b12e13f2ac88a8273d7bb2f6654e4750636aa122d6e853803e1b3b3ba

Observation ba276dbc-f849-4046-b967-aaad6c0b2f1f · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.256060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.256060Z digest=sha256:59c1a3984728818626c17249da6a05e1a0a27c2f3bc1ccd20b626f7f3c55bbf8

Observation 0220be99-caa1-4cc6-aafe-1062babe05ec · outbound

This paper cites A-JEPA: Joint-Embedding Predictive Architecture Can Listen.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A-JEPA: Joint-Embedding Predictive Architecture Can Listen

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.325028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.259155Z digest=sha256:4dc5e5ee268336f421ca29c8e38b3be9759da83f619d4ef1617ab64d933d8cde

Observation 4346d857-951e-48c1-a778-886ee9227d0c · outbound

This paper cites Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:47.136557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.261794Z digest=sha256:9832f7f2894b46ced3cc14e08b45a4c28eb1db0f1203c238b25ed1865c91c6d3

Observation f0292ae4-b77a-4c7d-af57-a5001062a6ac · outbound

This paper cites Masked Autoencoders that Listen.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Masked Autoencoders that Listen

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.264640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.264640Z digest=sha256:20358909eeb7ca4267d8a03a6f51eae3365a7eb641c5e8de0a88b95e0e3c8a4a

Observation c7a8b073-5e16-4d78-8d5c-3c06e9d4e34a · outbound

This paper cites A Dataset and Taxonomy for Urban Sound Research,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning A Dataset and Taxonomy for Urban Sound Research,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.267285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.267285Z digest=sha256:c31cf990e5ae5f0cb4f99be47aa652b7f7cce4770009058474ea6734a409958b

Observation 3f63ad7f-c74a-4ac1-9f15-6748b3153b20 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning VoxCeleb: a large-scale speaker identification dataset,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.270842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.270842Z digest=sha256:c1e19e3a96324b83d870b530d1d7964a8ed2a2ff2a98b4d5f2e0270e0d432e5b

Observation 2c3fbb49-5d1e-4fb5-8a87-4c07d3fb9534 · outbound

This paper cites The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.273343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.273343Z digest=sha256:eda05e8a09ea0d9ba5b1c39ad446a4a111265ba87ff51f411d82ccde8f0e768d

Observation 48c91fef-02c3-4ea6-b197-ebc56a608317 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.317711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.276048Z digest=sha256:144b84d74c3a4f399af575695a09196e80cb1cdeee9b96b044c3f42becd0d8d6

Observation 7e524576-c9a8-49e3-b7e4-fb065bb0e736 · outbound

This paper cites TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.989062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.278586Z digest=sha256:794dcb963c5a87cb84edbb11ee2e424e1478d3489fdb0b2546a753d18b930d0b

Observation 6463c55d-7c7d-4db6-932f-8ac0213d3ecf · outbound

This paper cites GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning GeoJEPA: Towards Eliminating Augmentation- and Sampling Bias in Multimodal Geospatial Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.976297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.282310Z digest=sha256:fd324d3bc3d6ce02823b84f1756a2d2f873d85842de898d072655aeff6e158ab

Observation 62a306d7-63fc-405c-8557-b8c1a04d2a8a · outbound

This paper cites Bootstrap your own latent: A new approach to self- supervised Learning.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Bootstrap your own latent: A new approach to self- supervised Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.309278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.284987Z digest=sha256:99468e930f070da30bd5d2e05a1991c25556d74c062268728811f46455691417

Observation a7d97c8e-cda5-4076-80b0-a38f88d9d71d · outbound

This paper cites Decoupled Weight Decay Regularization.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Decoupled Weight Decay Regularization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.287631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.287631Z digest=sha256:af209918769024e89ef21548668240102278f2c6c015b6eeead4e291f87fb3d9

Observation 43fa2dff-f79e-4711-a5f3-4176d3548695 · outbound

This paper cites an unresolved cited work.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:56:47.302159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.290010Z digest=sha256:9e3e47741cd493d0494fc2914832d2e656756e61a945c80c1c447ff8df521f8b

Observation 84372658-ed8d-4e29-b5b9-d66d883a3b6b · outbound

This paper cites Clotho: An Audio Captioning Dataset.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Clotho: An Audio Captioning Dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.295352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.295352Z digest=sha256:2266794fb68038b7dd770b9b49c383afa9f64a9bba0a03a3171622af0275fb14

Observation 0365b6ed-a8e6-48e8-a8ea-880de64dcfe8 · outbound

This paper cites CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.298919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.298919Z digest=sha256:6141a3b7f6d28683e5b88f3dd838bea871a90b66c5e73d309653740118f82911

Observation eaef854a-9b7d-40d9-9227-e85ca872837c · outbound

This paper cites Generating an item pool for translational social cognition research: methodology and initial validation,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Generating an item pool for translational social cognition research: methodology and initial validation,

Reference 26

Resolution
verified exact
doi, observed 2026-08-06T22:56:46.365463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.301331Z digest=sha256:3303111ffe1317b6e6650b6ff45b4a253b77f62d4d3f772d6eac8dd28ca50d2d

Observation 2c57c113-8cee-43b9-b57c-71e60fb42b58 · outbound

This paper cites Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.287063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.304258Z digest=sha256:1758c4aed597899c6159752282809229efa58bcae6c80aea4fcb174b2a89f3e9

Observation ad610024-4d77-49cf-a3ca-defdbfe393ad · outbound

This paper cites Sound event detection in synthetic domestic environments,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Sound event detection in synthetic domestic environments,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.279730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.306811Z digest=sha256:12ff9e038a07b31cf3279b496b27ec0ab2ccf41bf39b33f8bd3ac70065787327

Observation 971823ec-6826-4d3f-b24e-acbbf1fad5a3 · outbound

This paper cites ESC: Dataset for Environmental Sound Classification,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ESC: Dataset for Environmental Sound Classification,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.309460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.309460Z digest=sha256:8f5af8ed5b72bbb015f10d7640f3bc22a0615c7103064c34a83c31b806b2f230

Observation 9226004b-7c4c-4480-ba4e-5aa60fa31cbb · outbound

This paper cites FMA: A Dataset For Music Analysis.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FMA: A Dataset For Music Analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.311880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.311880Z digest=sha256:c1abaeb6f3cbe4cb26f699e0cda809105bf0e715e39431658ee2cfb26b9287da

Observation 6e457330-e699-4686-be90-3df5856309b2 · outbound

This paper cites General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.314564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.314564Z digest=sha256:9c43b794216d805b184d2f95bbc5b47b4d0b3766bb85061a0d35214cbffe5419

Observation bee0be84-1bd7-4e1e-a8d3-fcb2c2b3ce93 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning FSD50K: An Open Dataset of Human-Labeled Sound Events

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.317985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.317985Z digest=sha256:a79e4706952b8238ec2e4a2d3d061f67aa0ad6f767d0a0067114e33683e60bfb

Observation 2d5917d5-1fd9-460d-9a4a-134bdf7d73b6 · outbound

This paper cites LibriCount, a dataset for speaker count estimation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning LibriCount, a dataset for speaker count estimation

Reference 33

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.821091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.321292Z digest=sha256:87190e20c2aac9db5e6980f5ead4a3885e4e861768f0efc95a2a15b28e17d577

Observation e17815ff-6608-4292-9d68-a67f2e4f04ae · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Librispeech: An ASR corpus based on public domain audio books,

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:56:46.323549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.323549Z digest=sha256:1b4481047ea3b15ba5bd54b65dc5aa0deaff4f6555e978bc1d891d40cf161cc9

Observation f75daadd-0aeb-4571-a049-2a7d46df962b · outbound

This paper cites Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.327119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.327119Z digest=sha256:d62036b59a169158d956890478fea96fc146dbbd9122a7ac43f65660a477333c

Observation 25bd55ea-7fa6-40bd-8e6e-ed354dceb3b8 · outbound

This paper cites The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS).

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning The Ryerson Audio-Visual Database of Emotional Speech and Song (RA VDESS)

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.703751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.329589Z digest=sha256:d18949a5e5a0a9557f64ba0e9e1eb7493b704145c18042210a55a3a5fdf0f1c3

Observation 32864c85-9e8c-409f-8116-41152284d54f · outbound

This paper cites Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocal Imitation Set v1.1.3 : Thousands of vocal imitations of hundreds of sounds from the AudioSet ontology

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.623337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.332120Z digest=sha256:cf21a066c789f2687bd2e7cacdc8a0f3bd9553dd88dcc04ad1c05f5d60961ec9

Observation 614e6a51-cad4-40d4-b9bc-255abbffd4eb · outbound

This paper cites Vocalsound: A Dataset for Im- proving Human Vocal Sounds Recognition,.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Vocalsound: A Dataset for Im- proving Human Vocal Sounds Recognition,

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:56:46.334390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.334390Z digest=sha256:8f318bc015fb0ddbeb5bee2d2cc4720a677a051a1d5c001d799c1100456a84fc

Observation 26837dd3-70f9-4a23-89b5-4790b3d9fb14 · outbound

This paper cites voxlingua33 in WebDataset Format.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning voxlingua33 in WebDataset Format

Reference 39

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:56:46.486803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.336713Z digest=sha256:266666a3e587da93bf67fc2bd98c1a9f653f40c1839d4ed679e9b408fef70537

Observation ab31d341-a210-4f98-a294-fbf0fd3e9782 · outbound

This paper cites ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning ConvFormer: Plug-and-Play CNN-Style Transformers for Improving Medical Image Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.392405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.338990Z digest=sha256:4a306c38e7dc22d85d421322e7d1ac7dac38583770cae7b7fe4bf08cd7853e49

Observation dbd8b5b8-6aed-4331-a344-773037ffc781 · outbound

This paper cites MetaFormer Baselines for Vision.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning MetaFormer Baselines for Vision

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:56:46.381208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.342431Z digest=sha256:3de1b73126bef962936bee86ddaedd855904e836b142e400d339edaa3b010a12

Observation b4a2ae34-ed58-4a95-8c26-9d91d823260f · outbound

This paper cites Available: https://datashare.ed.ac.uk/handle/ 10283/853.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning Available: https://datashare.ed.ac.uk/handle/ 10283/853

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:56:47.294245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T22:56:46.292291Z digest=sha256:bcf6c95c9ba859fe1dcee5f6ea83f06402f5c80602d425f47da3d9e11641b596

Pith citing papers

Observation 72f1a689-e70f-43c9-9a42-817712dc2b54 · inbound

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning cites this paper.

Self-Distillation of Hidden Layers for Self-Supervised Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T18:11:35.194361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T18:11:35.194361Z digest=sha256:001beb544d6219241bdbd07490c24a563b29b4a2ee799774f4f5ee8e3c9814b5

Observation 27c28636-ee95-4fc2-9ef2-e1271c8ca576 · inbound

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space cites this paper.

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:06:23.815592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T12:54:49.511047Z digest=sha256:bda4d13505708a5906a060922c0233eee3ead82a4b6fc7fe00e624f86441a115

Observation d5760ae9-19fb-428a-aad5-0ae7642e0cc4 · inbound

Frequency-Aware Self-Supervised Music Representation Learning cites this paper.

Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:12.631693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T19:25:54.322923Z digest=sha256:e7416046e0b726ad6392041a66141a98949cd1f3ff15d118be1e8f3dab06a210

Observation 1bb2d02d-8af9-44e5-8425-e017f2a60551 · inbound

Frequency-Aware Self-Supervised Music Representation Learning cites this paper.

Frequency-Aware Self-Supervised Music Representation Learning Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-30T10:04:35.643472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T10:04:11.233115Z digest=sha256:c7ed7f07c51c38698aac01fdaa31c70f33b226ae98a61b7a6dbcf0f1297b13f5

Observation 8b7c4702-41a1-4256-b744-ec10d5d77758 · inbound

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition cites this paper.

Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T22:46:06.299864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:46:06.299864Z digest=sha256:0aea7bd837662a9152f2c6af9436f72c33b923159450329e657eb58f892d2645

Observation 163f1b05-700e-4916-8794-06d0bdd0842c · inbound

Music-JEPA: Learning a World Model of Sound from Action cites this paper.

Music-JEPA: Learning a World Model of Sound from Action Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:12:42.365990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:12:42.365990Z digest=sha256:a0cf6af0ed675f060ffed76489cb05f1d81f19254ef491a609c0259d9a24646e

Observation d67cc45a-d8e6-42ec-9e0e-eab62d207d31 · inbound

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds cites this paper.

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:22.852047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:40:22.852047Z digest=sha256:1c10781dcfd3176e479c01c7f6f7ece7269c423e7d136da920fb45459b4b780b

Observation 9e4fb628-aff5-4fb9-968a-da07888529a3 · inbound

FATE: Frame-Level Audio-Visual Temporal Embedding cites this paper.

FATE: Frame-Level Audio-Visual Temporal Embedding Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-15T15:15:06.622581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:15:06.622581Z digest=sha256:8f992516416e6f6e52141bcb25fb57ae83e56dd33d7c552dcc1253eec19f732e

Observation 1c325d45-a496-43b1-8277-e3b12ce316d7 · inbound

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture cites this paper.

Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:33.047406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:44:33.047406Z digest=sha256:cf3cc5f0290fd54b35667097f8515d91aef352b1b404eb77c1c83bbc0bb390c0