Pith. sign in

Paper Citation Record · LEDGER

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.23811.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23811 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-30T11:48:49.534136Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e21cd542-a294-44d8-934b-7df81d036d3f · outbound

This paper cites MusicLM: Generating Music From Text.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.415516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.415516Z digest=sha256:ec7821dd03b21f45d4b10c7d10fd1050086477c41e574fc5f3b6d9a46ab78328

Observation b2454572-4142-42e8-ad1e-0287e74a13f0 · outbound

This paper cites com/deepmind-media/gemma/gemma-2-report.pdf.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers com/deepmind-media/gemma/gemma-2-report.pdf

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.477289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.477289Z digest=sha256:732b7f4192bbcd9333c68aa52199a73d81be3deff833c2acc0a0f4da35386782

Observation 0b2fa850-bb71-4998-8e76-cabf201c68f4 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers AudioGen: Textually Guided Audio Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.491314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.491314Z digest=sha256:0e6c9b1662c470369e0f45c0d6153f4952efb962016533890e2ece1a9b0f3c35

Observation 2df537a7-7fdf-4fd9-b00b-816a65adfb4b · outbound

This paper cites Chunked Autoregressive GAN for Conditional Waveform Synthesis.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Chunked Autoregressive GAN for Conditional Waveform Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.506021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.506021Z digest=sha256:67d4188a5695c7d223de8228c2afe55eedeed0a675e2c656da825624e95424a3

Observation c6ed2877-ded8-4e06-8f7b-fca3dad09f39 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210,.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Librispeech: An asr corpus based on public domain audio books.2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.509573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.509573Z digest=sha256:2404381630cd174132445665f177b476ba6e639c324a010cd3e9a18deb9da496

Observation e00c8b65-13fc-4ce9-9780-4d7741398f39 · outbound

This paper cites Eval- uating speech features with the minimal-pair abx task: analysis of the classical mfc/plp pipeline.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Eval- uating speech features with the minimal-pair abx task: analysis of the classical mfc/plp pipeline

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.515999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.515999Z digest=sha256:ec5fb9e5450e0c4a7b3760dc936aa465af27e3e8c8642401ca8f2190ea5a5227

Observation 26515f14-db05-481a-9502-5459cc5e519f · outbound

This paper cites GLU Variants Improve Transformer.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers GLU Variants Improve Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.522615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.522615Z digest=sha256:e7f853dbdeab2693fd8c07ece53423e3983ce24e4e8d838e4d48201a1313b471

Observation e039e130-9fc9-4b33-afdd-ef05a5bfedf3 · outbound

This paper cites TS3-Codec: Transformer-Based Simple Streaming Single Codec.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers TS3-Codec: Transformer-Based Simple Streaming Single Codec

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.526693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.526693Z digest=sha256:63063db2f860880e6dd118db1db9515102f9e038b28f943861877f61aea5d74c

Observation 4592ddbd-ab3b-4b57-befd-05ec2ee93192 · outbound

This paper cites Biao Zhang and Rico Sennrich.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Biao Zhang and Rico Sennrich

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.530312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.530312Z digest=sha256:468eb3118c4aac80465ce537618d9ebacfe55bf51ea86270f4bc0d68973150f8

Observation 20638a92-3a74-4b07-9690-c58b1bf289b6 · outbound

This paper cites Root Mean Square Layer Normalization.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Root Mean Square Layer Normalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.534136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.534136Z digest=sha256:929ad20a175a080b6c1d7d02f73e256d446bd65aaa314caf8c9b083820bb5c0f

Observation 389770ee-ba4d-44f0-a81c-53148c577d9f · outbound

This paper cites William Peebles and Saining Xie.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers William Peebles and Saining Xie

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.512738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.512738Z digest=sha256:ec843172eb37781af041f2a8ae00fe9458790f0ff251c9c928cecc2c3b7e00e0

Observation c827519a-d409-4fad-8d7a-dbe151ca1f16 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.487060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.487060Z digest=sha256:0f159a6349861938976963dc138a94ada8d6ba6fbd7a0e8a45ab0a5252662ebe

Observation 6e732adf-c477-44e6-9729-136ae39cab44 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.502742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.502742Z digest=sha256:996291c3addc16dbc2ed92aed57d6faa5bf83277c586b0f5449c6906c6367436

Observation 5e9069ad-1b4b-4e21-b525-e7855bc98209 · outbound

This paper cites Longformer: The Long-Document Transformer.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.434595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.434595Z digest=sha256:9e484a8431a835399773d3ba6864a4d5a3f31cca0b31b965d48e4027677c5df9

Observation d0504c5b-bbea-47b6-a591-41777f8f3a14 · outbound

This paper cites Generative Spoken Language Modeling from Raw Audio.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Generative Spoken Language Modeling from Raw Audio

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.494905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.494905Z digest=sha256:e66ce0002b4edc6ea3cff25c517078ad50960504ba113d279800e33da18e48da

Observation 7db42bbe-1d03-45a6-8ca0-0af55c59bc70 · outbound

This paper cites High Fidelity Neural Audio Compression.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.461058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.461058Z digest=sha256:1a3268c2cd0e118b7767b7a68c79a4850767473bc920973acf67d1ee8581dc1f

Observation 221537d4-7c66-4dc0-9ef5-949e125132b2 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Moshi: a speech-text foundation model for real-time dialogue

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.443774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.443774Z digest=sha256:3660491ec644c1f07e8c2af3b2a12f4b72c098b4cbc508627294cb6be9ba3f88

Observation 7b841039-5eab-4ab0-9bfe-e88eaec1fdba · outbound

This paper cites Jukebox: A Generative Model for Music.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Jukebox: A Generative Model for Music

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.452676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.452676Z digest=sha256:e3782e508f34b9b8c3a6d6d59dbda756253e6c25ea3dccef40cc3837c03cbc45

Observation 9f2d315c-a744-44a2-b2e3-7246a5d8d1a6 · outbound

This paper cites Instruction-Following Pruning for Large Language Models.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Instruction-Following Pruning for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.482980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.482980Z digest=sha256:f3af6b5707ca919f6c5b6d007e28e0c6ec1bd5a19fc72f2246e7b59bdc233503

Observation 9a48891d-6705-4140-862d-2ffea9aa6f28 · outbound

This paper cites an unresolved cited work.

Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-07-30T11:48:49.469480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:48:49.469480Z digest=sha256:60441b4d5e2f6198a72933d62764fffbfa3bb273e137ba9801eca1e93ec77ed7

Pith citing papers

No inbound Pith citation observations are available.