Pith. sign in

Paper Citation Record · LEDGER

ETTA: Elucidating the Design Space of Text-to-Audio Models

As of 11 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2412.19351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19351 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:19.924941Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:29:47.889339Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T18:29:48.781610Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38a0f702-0c4e-4914-b2e9-c39311c56b32 · outbound

This paper cites write newline.

ETTA: Elucidating the Design Space of Text-to-Audio Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.665859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.665859Z digest=sha256:7421b345a6e60cb0fa2922acc7535bededc05ca47e1d9d42509b5a381bbcded5

Observation 0e949966-c858-4c17-bdc1-99c277bce3f6 · outbound

This paper cites MusicLM: Generating Music From Text.

ETTA: Elucidating the Design Space of Text-to-Audio Models MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.670861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.670861Z digest=sha256:719798d21d103043c8f51ae758ae906487aee599ddd31cd2599e4cb9737ee829

Observation 6bf6b427-dfef-4d06-bac4-0ccc9c0c2280 · outbound

This paper cites Improving image generation with better captions.

ETTA: Elucidating the Design Space of Text-to-Audio Models Improving image generation with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.675207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.675207Z digest=sha256:964ebc8b995444d01c9e76ea882c46156fd8fdbc9dc24bf0eb630cc3c0217ac4

Observation 79dc3aeb-1cd7-4c70-99b8-612df2b5bbb8 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.560211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.678777Z digest=sha256:3f3f1d7c0e5995f20bea591b2c761af1c51e37d48a135582f9e03e63f85fb179

Observation b26f449e-3485-4050-89b5-d163165cd51c · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

ETTA: Elucidating the Design Space of Text-to-Audio Models Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.682436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.682436Z digest=sha256:765bbc72baea16abaa3662a79788af3de61da13a9faf2a3662e8390a520497d3

Observation 6494ed4b-99c6-42fb-9346-7567e34056d1 · outbound

This paper cites Scaling instruction-finetuned language models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Scaling instruction-finetuned language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.686200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.686200Z digest=sha256:e2e25fba39c1c0d7c633f258593934f14e9a3364231546e1e808110fa6db58a7

Observation 81419a77-4f44-471c-a957-a01daf729252 · outbound

This paper cites Simple and controllable music generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Simple and controllable music generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.689887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.689887Z digest=sha256:4670bef81c28a1212ef99321797f347c67c518962b6d3ead9b1d885e069c5307

Observation 5318ae30-fa95-4658-8ecb-75886942d59d · outbound

This paper cites Look, listen, and learn more: Design choices for deep audio embeddings.

ETTA: Elucidating the Design Space of Text-to-Audio Models Look, listen, and learn more: Design choices for deep audio embeddings

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.535951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.693569Z digest=sha256:a5aef9d1bee6e067a9376b359261d38feb51ca99708b516c39b1becf962591a8

Observation 176eb31a-d2b6-43b9-8be9-8694f0afd531 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

ETTA: Elucidating the Design Space of Text-to-Audio Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.696800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.696800Z digest=sha256:379d8f8db1459756b64c1028ca885dfec7afe415614693bcb4895c8d254e0649

Observation 31aeb629-08cb-44a1-b4f6-94cf929500b3 · outbound

This paper cites High fidelity neural audio compression.

ETTA: Elucidating the Design Space of Text-to-Audio Models High fidelity neural audio compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.700431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.700431Z digest=sha256:dfdd406d32794fbe290a5313ebc63ab2679312fe5063eef327a1a653a254b640

Observation 9ea64442-a818-405c-a818-802cd5450134 · outbound

This paper cites Diffusion models beat gans on image synthesis.

ETTA: Elucidating the Design Space of Text-to-Audio Models Diffusion models beat gans on image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.703776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.703776Z digest=sha256:2ed017aa34ce648d8b370c361b61988d3e04c6ed0f864ff9abe1fb10b50015b4

Observation e0fe67b9-e09e-4f8e-bbc3-5ba7e3f01ba1 · outbound

This paper cites Natural Language Supervision for General-Purpose Audio Representations.

ETTA: Elucidating the Design Space of Text-to-Audio Models Natural Language Supervision for General-Purpose Audio Representations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.707054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.707054Z digest=sha256:b9d2935ac1c7370eb3b6f14c18e58060ffdd107fab1ada2e5d924320dd9da93d

Observation e2388ac5-ad3e-4afe-92c1-8884cc9c6897 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

ETTA: Elucidating the Design Space of Text-to-Audio Models Scaling rectified flow transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.710257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.710257Z digest=sha256:1aeabe0732bfc2eee6c1767f6022ec223d4051b230b009ae8f250c1c45f39b48

Observation e68c8769-783a-4b13-888c-6a9d48d3cb7d · outbound

This paper cites Fast timing-conditioned latent audio diffusion.

ETTA: Elucidating the Design Space of Text-to-Audio Models Fast timing-conditioned latent audio diffusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.503006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.712927Z digest=sha256:270a745e638473bf2e1447e78c64af5435e97991d7cbdd5c360cda2042f7f13d

Observation b4cf4c1b-afda-4ac7-bf37-dbac0f24533b · outbound

This paper cites Long-form music generation with latent diffusion.

ETTA: Elucidating the Design Space of Text-to-Audio Models Long-form music generation with latent diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.716105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.716105Z digest=sha256:d432fdc5914b34e04abf51d469e947f71d0ba0ded6a4c084886542079bd3fbf2

Observation 1bdeabe8-3bd5-4d4d-8216-902399c5c10b · outbound

This paper cites Stable Audio Open.

ETTA: Elucidating the Design Space of Text-to-Audio Models Stable Audio Open

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.719211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.719211Z digest=sha256:5501de8b53a816d10132aadc34f3f472435f7d5c62324bcc82531c51d44d8f66

Observation f7b5a11b-d485-4276-bed9-a1fd84d7200e · outbound

This paper cites FLUX that Plays Music.

ETTA: Elucidating the Design Space of Text-to-Audio Models FLUX that Plays Music

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.722159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.722159Z digest=sha256:f7b6431280967559d81d26c8f97465bd985d04959a6dcfab3fdc3b7d1441acfd

Observation 5ba8e9df-13c1-4f38-820a-23d50bc64b5e · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audio set: An ontology and human-labeled dataset for audio events

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.725729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.725729Z digest=sha256:8289ffa71c89eac0b0ab7326e6d9bdff44743988a74420c713f0befd172f3f2f

Observation febe2422-9f4d-443b-8096-6e1f6397f52f · outbound

This paper cites Text-to-audio generation using instruction guided latent diffusion model.

ETTA: Elucidating the Design Space of Text-to-Audio Models Text-to-audio generation using instruction guided latent diffusion model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.487304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.728509Z digest=sha256:1b52e61db3a354b5c29f50a9f6ec9f27694e462f0b16a1673cf0da810848d25f

Observation 19c8c6b2-6272-46a0-9a48-d7cf7a87aa43 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

ETTA: Elucidating the Design Space of Text-to-Audio Models EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.731205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.731205Z digest=sha256:132d7c67951c42db9d89b6a28f0bf44641ccdb89a5d43bea67b36a62008dc915

Observation 252aef73-a3a0-4225-925a-957d08ea0c46 · outbound

This paper cites Taming Data and Transformers for Audio Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Taming Data and Transformers for Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.734377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.734377Z digest=sha256:d22d71f715e64496e1814873fe4ac49f87d59cc1d186294fc56a7a3aba613c17

Observation 5d3b283e-dbf0-4bbb-bab5-91c692844f7c · outbound

This paper cites Efficient diffusion training via min-snr weighting strategy.

ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient diffusion training via min-snr weighting strategy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.478033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.737454Z digest=sha256:56bbe32d21739bd33e4191a2074739b254072329ff54433c653a2176a4d0c4fb

Observation ccebd0fd-8948-4f10-be94-e676e7ef8baf · outbound

This paper cites Gaussian Error Linear Units (GELUs).

ETTA: Elucidating the Design Space of Text-to-Audio Models Gaussian Error Linear Units (GELUs)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.740226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.740226Z digest=sha256:f525f25e4f4289d6afcd4ceffc61961f40dc7e5e8decf4858dfeb5b31f86c2c0

Observation b5f71e10-a9c7-467a-95ff-7ac8edaac624 · outbound

This paper cites Classifier-Free Diffusion Guidance.

ETTA: Elucidating the Design Space of Text-to-Audio Models Classifier-Free Diffusion Guidance

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.743569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.743569Z digest=sha256:b02e0588ca00b3905610db686c25f3482affb0c265b98a112880790017b8d493

Observation aa9cc778-8f86-43dc-bcdd-3e76505e4626 · outbound

This paper cites Denoising diffusion probabilistic models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Denoising diffusion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.747082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.747082Z digest=sha256:d21a4f788d3f1d5d3289196c6ac8c5a6355845c55871303ca0f049fff3c1592a

Observation cce53328-e6b5-416d-8390-c24d099da666 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.751403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.751403Z digest=sha256:de0980bdb0b49dd9058f86b36c5a09976aec4927d3e200184f274189b2852dcf

Observation 1ae4b6d1-0c62-45de-b91b-a1aa2d975d35 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.755042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.755042Z digest=sha256:5a0ae30bb72205376eb68ca17c9437de2854d3b04b0c278ae5ff4dca8d82a911

Observation 9d1d2ba8-4f67-4685-a5c4-a13af216205c · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.462914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.758831Z digest=sha256:f8e7377668f589ab93064d17f9ed3c5b9d15d9b31171341044bd3d0c98554903

Observation 9e6614b3-f58f-4196-894f-eba9dae087ce · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Elucidating the design space of diffusion-based generative models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.762087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.762087Z digest=sha256:8c9c3c88fd8062f703584c0097efc72ee5e68865076241aa1d02c069d432c041

Observation 0e0e4154-1a28-4081-ba3b-4e396ee4a461 · outbound

This paper cites Guiding a Diffusion Model with a Bad Version of Itself.

ETTA: Elucidating the Design Space of Text-to-Audio Models Guiding a Diffusion Model with a Bad Version of Itself

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.765403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.765403Z digest=sha256:3d3bf03a95f4a9b5628c5aabbac21825a092052b04ca46033040ef95d79e923d

Observation c41df97c-8c19-483b-b8f6-bf7d50b0b82f · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

ETTA: Elucidating the Design Space of Text-to-Audio Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.768891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.768891Z digest=sha256:fe6dbf3679c9117f89db1af45c79e594f7e65f43a5b26aafaff5886f35f96ff2

Observation ae9893b9-a990-452b-b519-6d748d3122ef · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audiocaps: Generating captions for audios in the wild

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.772358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.772358Z digest=sha256:3bbaea6510e0760cd9eeaa1d40b177c8b482f7aeb2e726965fa312cfc077983d

Observation c628301e-5e5f-40fe-b64d-e867819ac69a · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning.

ETTA: Elucidating the Design Space of Text-to-Audio Models Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.441346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.775405Z digest=sha256:bcacc76d69d61227795126dfb3be0a479b98feee649e3a8fcb7c0c6703b91576

Observation 6e596195-28a8-4d5c-a951-c396965664fb · outbound

This paper cites Variational diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Variational diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.778901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.778901Z digest=sha256:c09cbe5abf42004813cc428235a109995f8b0f06362608f7e0510ed8e5f5cc5c

Observation 751e8003-63d4-44d2-b6d6-8da49c09525a · outbound

This paper cites Kingma and Max Welling.

ETTA: Elucidating the Design Space of Text-to-Audio Models Kingma and Max Welling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.782110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.782110Z digest=sha256:2a019ad36a00d6ec92a0a64121ac18c95beaaed390f74ab4e8ba49c9c40f5838

Observation f2b717ad-6b0c-4f07-b1c9-d715030bd0c4 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

ETTA: Elucidating the Design Space of Text-to-Audio Models Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.420576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.785253Z digest=sha256:74b9b9d521996f1a0118de6a0f3a67ebcbfc5b67abd4dd6d1e3f89117bd10b34

Observation 4f4cc80e-6c47-4dd4-b606-13ce4387c1ed · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis.

ETTA: Elucidating the Design Space of Text-to-Audio Models Diffwave: A versatile diffusion model for audio synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.788372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.788372Z digest=sha256:9029e41de9620a1484210c01ac96536c2653e83784badfbe732ac245e1dbdf41

Observation 06599a77-24f1-49a4-8e98-00979612c6cb · outbound

This paper cites Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.404608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.791322Z digest=sha256:0ad93f630e1c86e544a53624925eacbf2f93c779819d9c68987e660900fac822

Observation b58e351c-68d9-4742-a8bd-d02e5e6ce317 · outbound

This paper cites Improving Text-To-Audio Models with Synthetic Captions.

ETTA: Elucidating the Design Space of Text-to-Audio Models Improving Text-To-Audio Models with Synthetic Captions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.794354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.794354Z digest=sha256:8d424301c1386269d22a53510c74b703486922853b8920a44f5c890af30bcba7

Observation 8c75c390-6775-4c6a-aa0b-9ae16fe73eed · outbound

This paper cites Efficient training of audio transformers with patchout.

ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient training of audio transformers with patchout

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.797269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.797269Z digest=sha256:677ada1d4e33f3e26c37e7e2d7a04b7d60f44d82e188395e7eb4fdc5ce91fdb5

Observation 40aa2056-1b35-4ee2-bf0d-1365d5d79202 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models AudioGen: Textually Guided Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.799997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.799997Z digest=sha256:6fef83669922ff7b474287d401e0f4490196c1187f8bfb8f9ecee65569326c0b

Observation 15833d4c-6bf7-41c5-a062-7be1a6f70a18 · outbound

This paper cites Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.803276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.803276Z digest=sha256:26876d5f18fb2845fb0364bf3c02cab3fa1cc7b870bd2b29061fbf08b40246fa

Observation 434b7642-ce13-487e-9721-c7a5fbee995b · outbound

This paper cites Efficient neural music generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Efficient neural music generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.395195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.806092Z digest=sha256:934e3786671cbd571711da0f429081be6dc9dac4995abee5a53631bc5a5685ac

Observation d3f9a6ff-0611-4c05-bf12-fc1470453737 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

ETTA: Elucidating the Design Space of Text-to-Audio Models High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.809107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.809107Z digest=sha256:93402ec419bfc64c605b5eb193ef9e14ed8c94bcaab143ff6854ff2926227de4

Observation c8ce5c63-f03b-446b-9ab1-a5b6fca078ee · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

ETTA: Elucidating the Design Space of Text-to-Audio Models Bigvgan: A universal neural vocoder with large-scale training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.812514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.812514Z digest=sha256:01207a84391164fa029505d676d62edafc51bb898817eb345399cc1c62efcccd

Observation 2e89695d-468a-438f-ac72-036c0e3b04b3 · outbound

This paper cites Quality-aware Masked Diffusion Transformer for Enhanced Music Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Quality-aware Masked Diffusion Transformer for Enhanced Music Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.815204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.815204Z digest=sha256:62f10132d7efabc37f7e453bc32509b2bcd4587b9a46a04dcc8a47cb431f9982

Observation 0309a060-3e82-41b2-b31b-7b2bd48de87e · outbound

This paper cites Jen-1: Text-guided universal music generation with omnidirectional diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Jen-1: Text-guided universal music generation with omnidirectional diffusion models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.818270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.818270Z digest=sha256:36e242143f1a98b8da9b8898e9b7749b31a4ebba3d2f4fdf591e2e02bf162e58

Observation 2631d82b-9bc2-4180-ad83-bca00b1c4835 · outbound

This paper cites Flow Matching for Generative Modeling.

ETTA: Elucidating the Design Space of Text-to-Audio Models Flow Matching for Generative Modeling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.821204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.821204Z digest=sha256:47fedee8374e2480f106e4d481a8252e90f18bd2be0bd513e089e42f66beb6d8

Observation eb2046cb-c949-4ddb-b41b-20cec4608d34 · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

ETTA: Elucidating the Design Space of Text-to-Audio Models Generative Pre-training for Speech with Flow Matching

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.824145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.824145Z digest=sha256:f00bc11f91c7edf2e2ebce8697d331c8325a9988d7348bb58b4f84b24fb4fc0d

Observation 4a28a237-5297-4e46-b9a9-8cea9619ae13 · outbound

This paper cites Audioldm: Text-to-audio generation with latent diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audioldm: Text-to-audio generation with latent diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.373890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.827269Z digest=sha256:f8a9865cff790b4886d84f446079e304d30ebd9acef48470717d63880e09fe8d

Observation 45117569-e7f0-4991-bf7d-81a43543f2ee · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audioldm 2: Learning holistic audio generation with self-supervised pretraining

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.830318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.830318Z digest=sha256:6a6b90eb0168df80cd29022510d93ad6fabb07f281c87630e3a70cf1fea24e6c

Observation 0f419579-be39-41dd-aa00-00ed85b49656 · outbound

This paper cites Decoupled Weight Decay Regularization.

ETTA: Elucidating the Design Space of Text-to-Audio Models Decoupled Weight Decay Regularization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.833498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.833498Z digest=sha256:6ccf4412426cdc7f6c937f454c960d22cebf62cd2a3056ba9b58c6754e1333ea

Observation 79d787dc-de0b-4ed0-b48f-810684e6a027 · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

ETTA: Elucidating the Design Space of Text-to-Audio Models Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.837772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.837772Z digest=sha256:601ad44572966f3a1038b1be5a671cebb276c08c2e685eb14a5c66d1a3a45ba1

Observation e98e2b40-e347-4d99-a31a-ac6cb5924f74 · outbound

This paper cites The song describer dataset: a corpus of audio captions for music-and-language evaluation.

ETTA: Elucidating the Design Space of Text-to-Audio Models The song describer dataset: a corpus of audio captions for music-and-language evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.359446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.841204Z digest=sha256:da76c5453b8eb41b9332e66d6879422666500f7c84e2a2e780af99ed653095b9

Observation 0ddf404a-7328-4679-9b8f-bcc20eacfdc9 · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

ETTA: Elucidating the Design Space of Text-to-Audio Models Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.844437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.844437Z digest=sha256:363cccbebc83acb2b0054bb934847e4939e28e34d2d8dcc080b3a4e6e5363d66

Observation bc78afad-bff9-43ca-be72-4027c2db333b · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Mustango: Toward Controllable Text-to-Music Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.847605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.847605Z digest=sha256:1ec60b895cfce7f348fe55578212e25bdd0d80ad096f8b3f3266a60e4e02521e

Observation c9baac49-05ac-4991-b9ee-b7bbb970e9a0 · outbound

This paper cites Mixed Precision Training.

ETTA: Elucidating the Design Space of Text-to-Audio Models Mixed Precision Training

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.851038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.851038Z digest=sha256:c402ebdf8bf5b2a329a9cde490adc7dbc8d076e064a23812e98f401ca43a16d5

Observation 877599ef-0ac9-4456-b80b-34f271828853 · outbound

This paper cites Improving multimodal datasets with image captioning.

ETTA: Elucidating the Design Space of Text-to-Audio Models Improving multimodal datasets with image captioning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.345071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.854997Z digest=sha256:4de5dbe5d3fbf9c588d5923b32e81035b8b3ca05983b9b5c25168890eb8c5bfe

Observation 15d34930-2554-46a9-b17d-d1ef2d63eac5 · outbound

This paper cites Gpt-4o: A powerful multimodal language model.

ETTA: Elucidating the Design Space of Text-to-Audio Models Gpt-4o: A powerful multimodal language model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.335939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.858283Z digest=sha256:d142c9d0f72d3f6cf9349d8ce330f08d833e61a00cb38a0499f370429d57d72a

Observation 4241cba8-4b8e-4b62-86f5-7adaaf966ad7 · outbound

This paper cites Scalable diffusion models with transformers.

ETTA: Elucidating the Design Space of Text-to-Audio Models Scalable diffusion models with transformers

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.861427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.861427Z digest=sha256:427bacc75fb6b95ef82dd70abfbf10f35c2c6be1730473a77da7187235f3b51b

Observation 5f5c2ea2-defe-4769-a2ad-5c8fb70ec2be · outbound

This paper cites Language models are unsupervised multitask learners.

ETTA: Elucidating the Design Space of Text-to-Audio Models Language models are unsupervised multitask learners

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.864621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.864621Z digest=sha256:472e11c1045a23868135897ff56222051450ed3e8b24c6d2b1c9f0e0746d2114

Observation f88baeb1-f7a2-4f7f-baf7-a6b3bdc47293 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ETTA: Elucidating the Design Space of Text-to-Audio Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.867788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.867788Z digest=sha256:867e30fc6642fd3f5fbb6c0d6f65e4a9db050e051b02f0cc2b73434fe5601e3d

Observation edb6bce3-9d89-404a-97da-5f7c813628e6 · outbound

This paper cites The musdb18 corpus for music separation.

ETTA: Elucidating the Design Space of Text-to-Audio Models The musdb18 corpus for music separation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.309155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.870873Z digest=sha256:9a365569acff00a52f04ce493e3d88008ec01e9ccf24350f2cf4b9c231d4deff

Observation 28c4716d-3ec0-4b15-b5aa-a296b03921c1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models High-resolution image synthesis with latent diffusion models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.873975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.873975Z digest=sha256:ff26b410ebc81c42ec60bc75762179e5c8face2b988d1a94b7dafcd58bd53f49

Observation b1f96203-c3cf-455f-b6a5-cdc7bcb28adc · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Progressive Distillation for Fast Sampling of Diffusion Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.877014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.877014Z digest=sha256:5633d31622a6a3bd87ecc0d480f79d8986818ec4bb5d5c29dd683206f7444f02

Observation b379d49e-be66-494c-a4a4-9a68ceeb91e4 · outbound

This paper cites Mo \^u sai: Efficient text-to-music diffusion models.

ETTA: Elucidating the Design Space of Text-to-Audio Models Mo \^u sai: Efficient text-to-music diffusion models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.292206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.880569Z digest=sha256:c4dc1a4733c598a6b19061bff3cc4045d463e8195ad8983f11a00a92e65a15c6

Observation 3010030e-b093-4861-8e03-ce84cb7aea29 · outbound

This paper cites Score-based generative modeling through stochastic differential equations.

ETTA: Elucidating the Design Space of Text-to-Audio Models Score-based generative modeling through stochastic differential equations

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.883342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.883342Z digest=sha256:6441c9eaa361e57fde4efd9c1c06dc51c2b3ba0c33857b05961b270942c0b25f

Observation 846d28ab-e886-49cc-bb39-4d6a2735d0ed · outbound

This paper cites auraloss: Audio focused loss functions in pytorch.

ETTA: Elucidating the Design Space of Text-to-Audio Models auraloss: Audio focused loss functions in pytorch

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.275969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.885821Z digest=sha256:4f932e191d33054893073e0c0418c348130bc1b8176bf20e3dcc2521f06303da

Observation 0160a295-6494-4ec1-8bb8-2b88f4d68bb1 · outbound

This paper cites Automatic multitrack mixing with a differentiable mixing console of neural audio effects.

ETTA: Elucidating the Design Space of Text-to-Audio Models Automatic multitrack mixing with a differentiable mixing console of neural audio effects

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:45:20.264696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-11T00:45:19.888546Z digest=sha256:04852aac586957819298cac0928f8bfdd730e525f4666a091e40e21607a702e6

Observation 7a763d9b-71cb-4685-a43b-a7803e2a0743 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

ETTA: Elucidating the Design Space of Text-to-Audio Models Roformer: Enhanced transformer with rotary position embedding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.891408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.891408Z digest=sha256:bed6a30743695bb21542b702167438d54a10073f2b370b7f1ef51059ba2bbfda

Observation 2ded199f-618f-4ce6-b8dd-9ee467885d8a · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

ETTA: Elucidating the Design Space of Text-to-Audio Models Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.894207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.894207Z digest=sha256:648d448341ae77eb88320d8988a282c02326d37b69dc509c528d2206b149a72c

Observation a28daef5-c711-4649-a51b-a06146b427f9 · outbound

This paper cites Attention is all you need.

ETTA: Elucidating the Design Space of Text-to-Audio Models Attention is all you need

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.897329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.897329Z digest=sha256:e45a1c24e7d915d2088028c88dfc0f2275ca146cacd9205dc29e79eab59c9b3a

Observation 72c87f90-3104-4ca3-aaaf-cd187fc35c2a · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.900293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.900293Z digest=sha256:c1220faf3fb9815c0b605ff683ac1b9627d908090438f60916c9e54865376d01

Observation 1b1ceba9-f1e3-476b-9189-44c649c5b49d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

ETTA: Elucidating the Design Space of Text-to-Audio Models CogVLM: Visual Expert for Pretrained Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.903369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.903369Z digest=sha256:be230bacacf0037d0dd200252e8746b39f3eec687ce6a6da9641c893887d161b

Observation ace7a19c-5d75-47b2-9b37-e27c2a52372a · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.906235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.906235Z digest=sha256:2d706504a016d8cc6f0a7e773abeed39710990b8f6f56784d8550818ed3f0f82

Observation 375f9341-8308-4fab-b4da-e3e22a06ffb4 · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

ETTA: Elucidating the Design Space of Text-to-Audio Models Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.908979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.908979Z digest=sha256:7a60ae40395b5e2e16d84b692f9b2ffe9330a782ffff1d89b69d5957e04b064e

Observation c81f82fa-c6b0-4067-9217-595cf16890a9 · outbound

This paper cites Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions.

ETTA: Elucidating the Design Space of Text-to-Audio Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.911892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.911892Z digest=sha256:194b2fc561f9fde22f274da1b0a92d33f3a3b4edc5ad2812329fc7f32b830cbd

Observation 43d9a63d-6bc0-4c65-b402-f1a61547b6ae · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

ETTA: Elucidating the Design Space of Text-to-Audio Models LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.914938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.914938Z digest=sha256:38e9a03eae8b4cd9fc09fd619913ef41c225c83d0fb408242bc45bc7a77fd87d

Observation 38af1539-d2ab-48a6-a29a-3a4df938e4e4 · outbound

This paper cites @esa (Ref.

ETTA: Elucidating the Design Space of Text-to-Audio Models @esa (Ref

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.917979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.917979Z digest=sha256:48d37c4840c3e007d20f5f5f3deed12ada1d507b8ccae95e7d1133bf7fbf3a64

Observation afe2ca9c-508c-489f-99db-a579f9037f64 · outbound

This paper cites an unresolved cited work.

ETTA: Elucidating the Design Space of Text-to-Audio Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.921516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.921516Z digest=sha256:f63eb3fc1761e022c6e28854375aaf75387f8d7c2af757194c5fe63e72ee8c3e

Observation 50b30830-190f-4d5d-8ae1-efca20564ce1 · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

ETTA: Elucidating the Design Space of Text-to-Audio Models LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.924941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.924941Z digest=sha256:0a47d57e6f9ce484b838bf9ab8cf174695ecf30e5c5aacc15591696c1f36c8ab

Pith citing papers

Observation 0c7febed-d781-4472-864a-180d04a118a1 · inbound

A2SB: Audio-to-Audio Schrodinger Bridges cites this paper.

A2SB: Audio-to-Audio Schrodinger Bridges ETTA: Elucidating the Design Space of Text-to-Audio Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:29:48.785905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:29:47.889339Z digest=sha256:68b9a4c02a015141f9a2cbbef96f0f51f6039a2f911f3d0f165505383e7fb0d7