Pith. sign in

Paper Citation Record · LEDGER

Masked Audio Generation using a Single Non-Autoregressive Transformer

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.04577.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.04577 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:54:23.104469Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.346881Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f090420-4e71-4c74-bd89-6aa68294f7f8 · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.513037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.513037Z digest=sha256:896fae2b557f06df2a514e8c8582b3a21fe07b67eb4c930b96bcded312192917

Observation 19a3a6ce-5f4d-4403-8478-40003504ead6 · inbound

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models cites this paper.

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:42.630260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:42.630260Z digest=sha256:cb4f51a49424b09223f4000096c3e57479c53cc0cd06b51a957a5094ba127c40

Observation 1cd8cabe-2215-4caa-af5b-c9d8f1144c54 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 159

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.215339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:487556e10e99bb7dc96f821e9bcb301a32423ae231c7b656f32bcdcd0c4d867d

Observation 794b50fd-0c54-4cfc-9538-8c3af8383193 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 118

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.348274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:1a4d6173a2b31e74594345b6e01cf5f4447afee701abb98382d9d555206db850

Observation b87d8294-457a-4af4-b9a7-7f18ddd3c5b4 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 118

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:97b87670c1e503ff4d51637b9da26c7e125d4d147bd9805f5eff7573b6d4ea93

Observation 663438e4-e295-4e3d-851e-943b5c9e488a · inbound

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model cites this paper.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.104469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.104469Z digest=sha256:2b8ee5177895344f63324d9de77598c602f18dedf7a859decd0d70f24a50d0c4