Pith. sign in

Paper Citation Record · LEDGER

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2306.00107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.00107 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:16:20.551474Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

27
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 648dad47-c0e2-4a1d-8b0e-2897f07b607f · inbound

SALMONN: Towards Generic Hearing Abilities for Large Language Models cites this paper.

SALMONN: Towards Generic Hearing Abilities for Large Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:29:46.289406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T02:29:46.242983Z digest=sha256:58ff78ead7ee2e31345aa1d4c59da331e05fcfac73d699f34c9091ebf463691b

Observation cbb8af77-3443-4945-bdf7-29efd5eff172 · inbound

Do Captioning Metrics Reflect Music Semantic Alignment? cites this paper.

Do Captioning Metrics Reflect Music Semantic Alignment? MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:16:20.551474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:16:20.551474Z digest=sha256:f6c723009d5dcfd3a23eb829f9d1642a7b46f767743255ff5eea17dd7677eac9

Observation b5be3d1a-7846-4d7f-a791-0cf5174197af · inbound

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models cites this paper.

MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T19:31:43.864877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:31:43.864877Z digest=sha256:cde16feb3325703bdb5d0ae29e240619840d28670b2b8746b0cf02969d8b8a6e

Observation 27369928-abb5-4a40-998e-8848e82acc7d · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.229355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.229355Z digest=sha256:06385209646cd799679cf3080b05389db710dddb25c718cb8ad5586cd9a9dde9

Observation 6fa4d1ab-61c2-4b4f-b250-73617929a07e · inbound

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization cites this paper.

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:22.065862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:22.065862Z digest=sha256:1966e49ac9039e769895845ce5d400b7f4af65ebab7c2114d811e4f420137d1b

Observation 8653f544-404e-42e8-8a25-cb98b346e98b · inbound

Towards Unified Music Emotion Recognition across Dimensional and Categorical Models cites this paper.

Towards Unified Music Emotion Recognition across Dimensional and Categorical Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T00:06:08.261142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:06:08.261142Z digest=sha256:0c8cafad48ff5827779287f36f6f83131b741de46929e63e09c2bdd9dd5270c7

Observation 6c9ae4b0-9bca-41e5-a443-ecfa11e4a866 · inbound

Semantic-Aware Interpretable Multimodal Music Auto-Tagging cites this paper.

Semantic-Aware Interpretable Multimodal Music Auto-Tagging MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:51.845493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:51.845493Z digest=sha256:849696397d64c65cf7e2e083547edb3a9640fed90d249c97631dfceef73312e1

Observation d38b7e93-573d-4b25-b831-06d3394bab70 · inbound

Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction cites this paper.

Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:29.869542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:29.869542Z digest=sha256:0fadc9c4cdcb1624b2b5a71c73a938c04b89cc74077266341b4fbd28d01b8de6

Observation e4069a85-3333-437f-b636-a2c8f78ba737 · inbound

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models cites this paper.

Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.894022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.894022Z digest=sha256:10813d27e081a5b41f4e23a0e88a55c9a083005d21823a3cb607316bb544e3f2

Observation 174d3af3-f042-43d5-b345-337e0e8e60ae · inbound

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding cites this paper.

DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:58.255447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:58.255447Z digest=sha256:be2aeec94d18d828fc968a21b43ebd45f895df86e9dccd0e5cd66ca62101d3df

Observation 7af0fa36-0256-4ed7-9d7d-a9600c00945e · inbound

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance cites this paper.

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:16.403568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:16.403568Z digest=sha256:b2b95501f5a5a55d33f859e8ffc349f47848d64bcc867b763efa9e14bd28dbee

Observation 515c194d-e8bf-4af2-a06d-42134746d7af · inbound

Workflow-Based Evaluation of Music Generation Systems cites this paper.

Workflow-Based Evaluation of Music Generation Systems MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:03.252951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:46:03.252951Z digest=sha256:bed9c4b6de7d11551eb81ec638a498a4d37b3393e1647abb0a6b2b9a81020265

Observation 094ec0fb-8f30-4326-9ff4-c0b2f93231c3 · inbound

OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction cites this paper.

OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:10.687382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:14:10.687382Z digest=sha256:2aa1dc3b59fbe6b90832a053316389cba64bd545d468c10e64b16b8aead6aaf5

Observation 6a77f882-56df-428e-a1ab-501cd04a7761 · inbound

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation cites this paper.

MusiScene: Leveraging MU-LLaMA for Scene Imagination and Enhanced Video Background Music Generation MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:20:44.402435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:20:44.402435Z digest=sha256:b58ac1c54c729e887270ffe46f34455a664b76a4b6c2a8962073c93fb807dc62

Observation d3551d87-3533-466c-8827-07c207307ded · inbound

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing cites this paper.

MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:03.277188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:12:03.277188Z digest=sha256:d0f7af60cecb569fdf1b9021a71216cfccccee7f50fee28b1cf518bbf7a2cbaf

Observation 8bb2fca9-b7d9-4b04-8130-381de20ba017 · inbound

Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models cites this paper.

Enhancing In-Domain and Out-Domain EmoFake Detection via Cooperative Multilingual Speech Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:30.038796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:30.038796Z digest=sha256:846e5bf4132648fb89afb4630190d8f7d2e26d65392b483e3b71e159b122a1d4

Observation ab597606-0dac-4cf4-9ba7-620af391b0f4 · inbound

Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation cites this paper.

Affect-aware Cross-Domain Recommendation for Art Therapy via Music Preference Elicitation MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:21:59.744231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:21:59.744231Z digest=sha256:bc90d601ed59d09868308532e51de272ac1b660f3db1729076b9e4c927553a85

Observation 3e291965-80a8-4978-8dbd-981f289eabac · inbound

Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack cites this paper.

Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T05:48:41.966906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:48:41.966906Z digest=sha256:de785962bd545adcb7acc1a0e7340e33c11fc4ae86aef88e2d198835bb0136c8

Observation 95a1f3cb-8e23-4042-a4e2-af70ee365a5e · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.781060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.781060Z digest=sha256:ddbaf4b5b0d0ad30e9db571d89231e057448d245fd5ead476c5c3c40f13cfb3c

Observation 1aff77a5-2bc7-4da2-9f0f-ec9d7ff429c2 · inbound

Exploring How Audio Effects Alter Emotion with Foundation Models cites this paper.

Exploring How Audio Effects Alter Emotion with Foundation Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:34:52.424527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T12:34:31.239629Z digest=sha256:5ac4d6b20580ef16421bf82737f984fd03c97b1ca4ea665520f1c74d6a1a079d

Observation 466a9129-a629-42a0-99e4-52db9dc7a8e6 · inbound

Assessing Factual Music Comprehension in Large Audio Language Models cites this paper.

Assessing Factual Music Comprehension in Large Audio Language Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:29:46.662633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:29:46.662633Z digest=sha256:640e57996b46ac27bc3a75ccd978dd9512780f79cecb14a7ff192ba8a2b63c12

Observation 7a0a1d41-2097-49f4-9dcc-a14f3e366eb4 · inbound

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity cites this paper.

Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T11:40:03.186955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T11:36:06.967549Z digest=sha256:f67585aaa4a5b8f6e9eea1a1f8cbbf8c34e0c73ca37db50739114f78d396726c

Observation 5fad3ff9-fb68-4ccc-a8bb-054687c776bd · inbound

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis cites this paper.

Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T17:17:10.597653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:17:10.597653Z digest=sha256:5ebec83b680e199fa056113418949286cab4aebdd1b93051bce8009b8be94e62

Observation cce0c1e3-c4fe-4243-b464-6a55876b087a · inbound

ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics cites this paper.

ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:26:59.772765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T07:22:07.949007Z digest=sha256:67d31fba853d851acdc21eedfd0657c80b8bf382cbece707b8717ff1dcd84b8b

Observation 9c5f152f-f46a-4ba5-857a-118d3f5f3fb7 · inbound

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models cites this paper.

Revisiting Content-Based Music Recommendation: Efficient Feature Aggregation from Large-Scale Music Models MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:40:30.796961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T02:39:46.434831Z digest=sha256:7d7a1e1331611ab6d450d15c8215270a6317e872c73a40a8f2d4bedee305c303

Observation 82206991-8955-4217-90b9-8f2ba3d275c4 · inbound

Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems cites this paper.

Adopting State-of-the-Art Pretrained Audio Representations for Music Recommender Systems MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:01:13.325918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T07:17:42.922029Z digest=sha256:877146100805340c9ff3047d1a6125cb473f7d1d436dea3ed611c2305f49a727

Observation afb7bdbf-5781-4f19-8bd7-74eddb440120 · inbound

ARIA: A Diagnostic Framework for Music Training Data Attribution cites this paper.

ARIA: A Diagnostic Framework for Music Training Data Attribution MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:27:43.044419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T18:24:29.290904Z digest=sha256:cd32be2c74fb35807389b883803522f53b65ebf2403546a77fbc2fe8227935b1

Observation e292391c-f19b-4486-a2fb-a0b7211ec3ad · inbound

MERIT: Learning Disentangled Music Representations for Audio Similarity cites this paper.

MERIT: Learning Disentangled Music Representations for Audio Similarity MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:43:32.564872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T15:40:57.719214Z digest=sha256:5bee58b28ad93a9cd75ed5948435d3ab2a61672bb284e07d22a95ab4f5548519

Observation e2d95631-def2-45f2-b5ff-21ce2edf6bf4 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.655976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:6eb1889821e34abd8f8c6572af2ffa088d1df10deb5aae63689f374da9bb3d65

Observation 5fbd5ea8-fb63-4e40-b732-416492efbf74 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.890564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:946a68205b0620dc3c7c876fbda68a85c3321ba83955ceb94b4251f226c2dc9f

Observation 6dfb6e8a-eceb-49a8-86e2-e512d2929b61 · inbound

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations cites this paper.

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T22:56:37.693891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-09T22:50:19.910504Z digest=sha256:9b82567bf0bc9df03bdd1576164f0a166dd32e91d77483460221f95486fb3f78

Observation d83fa3d8-300e-4e25-b351-e4031f6ae035 · inbound

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment cites this paper.

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T10:59:22.914974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:59:22.914974Z digest=sha256:8455911be355979fab2370cd08438b76d906775115731e063fdf103a1273d470

Observation b1f2f3da-f1e6-48bc-88ac-3a7e80071f5a · inbound

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems cites this paper.

StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T22:48:42.353079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:48:42.353079Z digest=sha256:ee626da6fba45340fe0713bf0ea7e2fd9bd8de69c64c881ef3aa83d70104f96d

Observation c819db37-9690-4c80-8a36-6ab584be0a12 · inbound

Do Music Foundation Models Embed Pitch in Helical Structure? cites this paper.

Do Music Foundation Models Embed Pitch in Helical Structure? MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T13:57:15.866365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:57:15.866365Z digest=sha256:d7ecf1223b0104dde7c4b0870382297628d00130fe8bb4bd4bd91b3311f2d107

Observation 5edca3c8-d14f-49dd-a1f7-e91746e43062 · inbound

CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control cites this paper.

CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:38.125621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:38.125621Z digest=sha256:06d21b103142119517aff5a540da62693f25cca28e19eeddd5c2cbefe5d9d13e