Pith. sign in

Paper Citation Record · LEDGER

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.04378.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04378 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:42:37.133267Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved15
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a8874bf-1f7d-4649-9a84-57fa7402b568 · outbound

This paper cites Self-supervised learning from images with a joint- embedding predictive architecture.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Self-supervised learning from images with a joint- embedding predictive architecture

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.080911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.912135Z digest=sha256:0393b9d66fca252ce5a940735f78a1e0da1cafb8e7b9c4695e43213f9050c142

Observation 87ab2416-fe04-4adf-a606-c7d2920dc467 · outbound

This paper cites LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation LeJEPA: Provable and scalable self-supervised learning without the heuristics, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.063959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.918353Z digest=sha256:4d1fa1be2338a507c412aa2c0d41bf3e4836d20a5649b32f6d5dbfd65d6f5eec

Observation f0ce2832-a2bc-4771-a2d8-30745e4a1e7d · outbound

This paper cites Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Bittner, Juan José Bosch, David Rubinstein, Gabriel Meseguer-Brocal, and Sebastian Ewert

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.048869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.923728Z digest=sha256:f46f081aecb7b87e569f1b93b95ad5cc28f94845ed1d6aa9488f7afbc9c7034c

Observation 2cf73a05-3977-49d2-80f3-fc2571c3f015 · outbound

This paper cites Zapata, and Xavier Serra.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Zapata, and Xavier Serra

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.032439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.929545Z digest=sha256:d9fc2850088733b546e38e89f67eb7420390364753b301abf4f3dda7b05865c3

Observation 8ab0eeaf-e0d4-4b3c-90ec-aca0412ce320 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Emerging properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.017269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.934794Z digest=sha256:0ee64bb2e0a54298902ba774e201ddf7ecb5855270e78632219da5c5c2845c0a

Observation 59777f34-56c3-4d60-b615-a86b8be63621 · outbound

This paper cites Codified audio language modeling learns useful representations for music information retrieval.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Codified audio language modeling learns useful representations for music information retrieval

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:38.001701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.939949Z digest=sha256:a9dcd14818414cd5f28670c2ebd24e64f5b170368fa4354076ba67076ac167f2

Observation 5c3baffb-2a94-4a1e-8893-96df08d80ae6 · outbound

This paper cites Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.568953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.945515Z digest=sha256:69468a1957b2a431f8834bbd448e0d73e2013b678f4467210d605f0a57a3eebf

Observation 8b25a6df-49bf-4ad7-9374-fec2b95402df · outbound

This paper cites Pixelflow: Pixel-space generative models with flow, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pixelflow: Pixel-space generative models with flow, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.950691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.950691Z digest=sha256:ad6aa382705cca1e6cc3e6b8bf86135f1534f8c5a2176ada86a1e53e0aefd2d7

Observation b76b3215-d6ea-4ee0-a166-f468f897a65c · outbound

This paper cites Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Automatic Analysis and Influence of Hierarchical Structure on Melody, Rhythm and Harmony in Popular Music

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.955570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.955570Z digest=sha256:ae725b2866592364cc6a69142dcfa2043436b403b7f8883e49c7d7bb506f3786

Observation 0872e412-f945-466a-b80e-e54c31e0fc54 · outbound

This paper cites Diffusion models beat GANs on image synthesis.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion models beat GANs on image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.976481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.960678Z digest=sha256:1504212c81aaa956311ddaa75a89a3f4f4de4407c3a444894f92668314d7f9f7

Observation 99673cb7-a190-4952-9601-3eaf6166e80b · outbound

This paper cites Generative modelling in latent space, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Generative modelling in latent space, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.962049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.965421Z digest=sha256:8d83495b9fc713373d0b0db7ea0a7abda1cfcd4a9a1af230f336da53f6670f56

Observation ee53d3be-8918-411f-a66a-0973ca62549c · outbound

This paper cites Hawley, and Jordi Pons.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Hawley, and Jordi Pons

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.946586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.970418Z digest=sha256:a531ba078f53ea836d4f69f0e4fe5c12070d9b5c265178188931b58646f99a3f

Observation ec1dd9e5-264d-4417-b06f-3f8a9553fca9 · outbound

This paper cites Learning and Leveraging World Models in Visual Representation Learning.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Learning and Leveraging World Models in Visual Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.975241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.975241Z digest=sha256:c372b85788b17e5bb50f89af34ebd8be8b4d6790c54aab7b3bde6c3b8662a3a4

Observation 823ffb86-9a19-46ce-8fac-aed9cf3d4be2 · outbound

This paper cites World Models.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:36.980422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:36.980422Z digest=sha256:1d7412a6f8120bd44355f5c966dd3d1f3711f2434129344e01d131df2d22499d

Observation 19e4507f-3ca4-4b83-985a-1df8eaaef8b7 · outbound

This paper cites Using a joint-embedding predictive architecture for symbolic music understanding.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Using a joint-embedding predictive architecture for symbolic music understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.929987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.985406Z digest=sha256:9839bb7dae73621c7777f05cceba58373e3b169b05a67cf24e9e07959ed70b7e

Observation 2a4bf12f-9026-413b-a7da-0e8174363d51 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.914256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.990376Z digest=sha256:c4453c646fc27cd1c20451961f05ade1df6f13f8bd6eea2e0c36c7aac0068846

Observation dcf967d7-5d7b-4c45-9c3f-0b29962cccbf · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.898382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:36.995429Z digest=sha256:47707cd8592e9e8c6db43ac7cddc24eac5214126404ad1ba951a7ec22e2671db

Observation 58bff056-ca19-4e5f-94e5-a646e70108bd · outbound

This paper cites EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.882781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.000539Z digest=sha256:2419198d52761f1fb08655996b6d53130eb1095224d8a2631e014a416dc319d7

Observation 776a0ebf-080a-4d0b-bcc9-900b302cffdc · outbound

This paper cites How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation How far can pretrained LLMs go in symbolic music? controlled comparisons of supervised and preference- based adaptation.arXiv preprint arXiv:2601.22764, 2026

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.005741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.005741Z digest=sha256:71aa00d8731d5334066eb52bcffabd1e6e4e22cda40fa946ad0229dc1d069aa5

Observation cb009084-0471-46a6-918f-b6819232b411 · outbound

This paper cites A path towards autonomous machine intelligence.OpenReview preprint, 2022.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation A path towards autonomous machine intelligence.OpenReview preprint, 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.867340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.010781Z digest=sha256:c72fb3dcd1368bfe2fea104a674209a38807b1c1296a884277e91e49cbee6800

Observation df7df0f6-71f8-4b7f-92f3-da69e35974af · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.015774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.015774Z digest=sha256:d2d1d861b532c30b47e656344712454c7337f01d01d2f9bee83fe6bb75438dd5

Observation ed68fb50-386a-41f0-9356-2fe0385ab834 · outbound

This paper cites Swin transformer V2: Scaling up capacity and resolution.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer V2: Scaling up capacity and resolution

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.841633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.020732Z digest=sha256:a4ea24b80f4e49b8b65f5df0a3091560e1107b06f9e32248311595b4235b98c6

Observation 23e7dc7b-507c-4dab-86b3-6699f2f59d6f · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.826192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.026096Z digest=sha256:bc89c3babda3ed3af5bc184d38add2f11203054eb1afe377b619e3ca1fd8eac0

Observation f4fec9ca-7df3-44f8-9754-a72c49232f45 · outbound

This paper cites One-step latent-free image generation with pixel mean flows, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation One-step latent-free image generation with pixel mean flows, 2026

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.809553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.031151Z digest=sha256:004b7529c76a4fcf29a292f0a6b7aae521b2a9f071c0c6782d64c246ab444081

Observation 2d36acc3-76b4-4d96-a9b7-a0d8875d630b · outbound

This paper cites CMI-Bench: A comprehensive benchmark for evaluating music instruction following.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation CMI-Bench: A comprehensive benchmark for evaluating music instruction following

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.794140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.036013Z digest=sha256:1590bed9e53d5c1f5be5fd65967ed1911c3ac1a21372db6da99cef650f6c5af9

Observation 75e79708-a7b9-4776-8996-3be58b18a4d3 · outbound

This paper cites Pnp-flow: Plug-and-play image restoration with flow matching.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Pnp-flow: Plug-and-play image restoration with flow matching

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.778168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.040895Z digest=sha256:45434fa71392f9d3fa01566fde0b567045017030181837aeb2de9bff1fb4182c

Observation 85c8f97e-4f54-4634-8f48-c8cedf74184f · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.762821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.045834Z digest=sha256:d0080b7f48c6197110b65e8b876b0bade30c962cd2bdb3985258120c7adb2dbb

Observation 134e1ddc-edaa-47a2-b3c8-48fd902bb68c · outbound

This paper cites Polyffusion: A diffusion model for polyphonic score generation with internal and external controls.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Polyffusion: A diffusion model for polyphonic score generation with internal and external controls

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.747648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.050886Z digest=sha256:121496308ba8671e6e85832e7f11204410ccb8115a2d7b94840955362af47c82

Observation c8146f13-b4f8-4ba4-a32e-9f4b3cf33c60 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.055751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.055751Z digest=sha256:d75f72634558a1c42b07fc5e99111825c73c8cb86af540b7e4dfec83950f6f43

Observation cc2dee26-7156-484b-99a7-ad1273520a00 · outbound

This paper cites Muckley, Ricky T.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muckley, Ricky T

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.060824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.060824Z digest=sha256:d75ecba41778eb531121d4a87f4e401e0a88fa593582b15e7c59086ccd474a9a

Observation 40c9a460-054b-4def-828c-bc6a56b1a33e · outbound

This paper cites PhD thesis, Columbia University, 2016.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation PhD thesis, Columbia University, 2016

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.720976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.065674Z digest=sha256:acb4b62dc8b7e8f02b6bbf1c6c152ff271ceaf6435053f04c8bb709ac77fede1

Observation 223d6f9e-ba15-4769-9dcd-69218455e90f · outbound

This paper cites Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Stem- JEPA: A joint-embedding predictive architecture for musical stem compatibility estimation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.703918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.070460Z digest=sha256:af9b9dc071126542dfd68fc49e6b5a666debe8607aae99b329be14adf1ed6ccb

Observation 0f15c106-a328-41bd-89bd-081b0e7f38c1 · outbound

This paper cites Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Crash: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.688676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.075034Z digest=sha256:9a8ae5d479a063de2dbcde02c1279380670f1e16f589a309d453694ace0f4862

Observation 6bb44a12-d538-46b4-8001-cd46248a1f9f · outbound

This paper cites Muscriptor: An open model for multi-instrument music transcription, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Muscriptor: An open model for multi-instrument music transcription, 2026

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.673242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.079711Z digest=sha256:9577fe1639d9ec80492210e364d8e18ac7608f4bffbb2df12079aadf0712f1e7

Observation c226ce58-bccb-4284-878e-3145fae4044d · outbound

This paper cites Rick Rubin: The 60 minutes interview.60 Minutes, CBS News.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Rick Rubin: The 60 minutes interview.60 Minutes, CBS News

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.655671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.084436Z digest=sha256:b29c67051bd61795bffaa28e309ac2ad1425bcd4d254ca323623ad8b246ac3fc

Observation 2c5f8fec-afaa-4ae4-a232-96c8a90cf810 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:37.640139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.089482Z digest=sha256:5a14668d89cba02bbe2e62134628fe387513ac66b87a6baea8ed9c1cef0ddac1

Observation 36cb00f1-254e-4088-8729-14943ecbb7c7 · outbound

This paper cites Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Improving and generalizing flow-based genera- tive models with minibatch optimal transport.Transactions on Machine Learning Research,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.093962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.093962Z digest=sha256:b1eeaa57ae83327f0f19066a6719ebc3a9f38e588139a03dbcf3ab763cf4db38

Observation ab69694a-1cbf-45f4-b5c5-310b296da0c0 · outbound

This paper cites Music-JEPA: Learning a world model of sound from action, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Music-JEPA: Learning a world model of sound from action, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.601766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.104016Z digest=sha256:38a81fa4742ea22c53849bf5b64c95881d7f64c559f77c4a7f813158b44677c8

Observation ef97b942-4dd2-4979-ba95-68a7039b397d · outbound

This paper cites Visreg: Variance-invariance-sketching regularization for jepa training, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Visreg: Variance-invariance-sketching regularization for jepa training, 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:37.585245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.108727Z digest=sha256:b7a0fefb8e9b14508017e65ae77ebee6b016ca28a1148ba0c4399016b89eaab3

Observation d4332ef1-0b75-4283-a70f-c45e98a78200 · outbound

This paper cites MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MelodyGLM: Multi-task Pre-training for Symbolic Melody Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.410262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.113557Z digest=sha256:db4db562d7c3f1df1e48231b24da305f463a558628a11316e906c2610c6f4235

Observation 8564d54c-c355-4bf7-ade4-163d2526236d · outbound

This paper cites MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation MIDI-LLaMA: An instruction-following multimodal LLM for symbolic music understanding.arXiv preprint arXiv:2601.21740, 2026

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-08T18:42:37.387019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.118857Z digest=sha256:59bf37b1a7f03730200bb47009d50d4d4fcff967bba8b8032227b8d8337d4d51

Observation b90d367e-3164-4d36-a123-753077ee38f0 · outbound

This paper cites ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:37.310363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T18:42:37.123392Z digest=sha256:a097b96e5203de2d006835db12681949dd4b7d23ed2cf5773d1856ae0f51575d

Observation d31e8fa0-d1c7-49c7-bad9-438cbb904560 · outbound

This paper cites ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation ABC-Eval: Benchmarking large language models on symbolic music understanding and instruction following.arXiv preprint arXiv:2509.23350, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.128449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.128449Z digest=sha256:6cebec90fd2abebfec421a52be0e62cb92514c84cf0e42b1a4d6b7492896bf1c

Observation 9c36c57d-20fb-411d-b5cc-6c78bc1b0432 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Diffusion Transformers with Representation Autoencoders

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:37.133267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.133267Z digest=sha256:fee6b2555ad873a7af6e263ce24f8f098a9bc0eb89f8d5d47d3e52c9f758f554

Observation c4c16f69-7feb-45d3-b797-efdfae11cf75 · outbound

This paper cites an unresolved cited work.

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-08T18:42:37.099100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:42:37.099100Z digest=sha256:f6962f045b1a026343d72a779b52d8817a2203ac772926d4bb924111e24139dd

Pith citing papers

No inbound Pith citation observations are available.