Pith. sign in

Paper Citation Record · LEDGER

VampNet: Music Generation via Masked Acoustic Token Modeling

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2307.04686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.04686 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:22:17.883184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T11:45:47.216407Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 214f2509-3365-4596-b4c2-900f898a8b6c · inbound

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models cites this paper.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.883184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.883184Z digest=sha256:8b3bf87d5732fa0d2c3b495a51c2f188bc658afe2300d421c4024f000d56a514

Observation c64f2dcd-6e0b-43d4-8aae-b5b26ed4d687 · inbound

Video-Guided Foley Sound Generation with Multimodal Controls cites this paper.

Video-Guided Foley Sound Generation with Multimodal Controls VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:00:01.025959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:00:01.025959Z digest=sha256:b94b61b341647b0e0c7ab2b9211945ffa9030997e1d4ee800975ea3d7e6ff783

Observation 79c9ba49-fd6c-4aa0-9051-3c09f957c784 · inbound

Watermarking Training Data of Music Generation Models cites this paper.

Watermarking Training Data of Music Generation Models VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:49:42.555192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:49:42.555192Z digest=sha256:4ce636f119e9cdc15403a2f1b54f06bebc2bb399ab7b2c7154d30ce9b16450f8

Observation 6cbe2a95-0608-4a2e-a561-090396b63775 · inbound

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor cites this paper.

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:52:03.761841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:52:03.761841Z digest=sha256:ab026b479c0549a3d00b59438c91c7b40cb2779921a86851065180477997dcb8

Observation dd5d1d1f-32ae-4e81-ba9c-a87c315d67b9 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:24.581035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:24.581035Z digest=sha256:a496e5b2fd79fb610edfa6b6c3519b241a276e3099144823d5f6eb6e8ce5e9c9

Observation 5cee1096-fb04-4952-b481-c1f2c712c28f · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:01.492724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:08:58.648967Z digest=sha256:5822673560894ad286ddb4cad00a4451e0b00b96665e20754fa804670093e3ae

Observation e9fef923-e604-4b24-b57e-e3d382c359cb · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:21.106772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:21.106772Z digest=sha256:7c2ae9ea2343827e0fe173fb2bd201efdcf8d7aa933cea040d122eb08c7b5600

Observation 1f8668c3-a2cf-4d3b-8513-92dc331adfbb · inbound

Latent Fourier Transform cites this paper.

Latent Fourier Transform VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:06.999351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T03:45:07.892234Z digest=sha256:0b86c756fd9ec507f61fc58e67a4bf93546e3017faf9a34327bdf94195ae0433

Observation fd56e338-9c70-4e3b-b60e-b98286b6bad7 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.218495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:3295aa33187a8ef92cc1a7d56431c2f14bfb41c3fb645f3a7ab5f5fdca858db1

Observation ddad399b-a31e-44e0-9fdd-41e2f35af40f · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model VampNet: Music Generation via Masked Acoustic Token Modeling

Reference 158

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.218036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:cce2923fb81d95151172161dd04e2b0fb1bf29ad7aef000844698a83dd7b6103