Pith. sign in

Paper Citation Record · LEDGER

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2412.09789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09789 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:48:46.626775Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T14:53:15.718359Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T14:53:23.238659Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95bad160-aa81-4f78-b5ff-c140d91d4d23 · outbound

This paper cites Stable Audio Open.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Stable Audio Open

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.503987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.503987Z digest=sha256:93251ca9c0cc088efbb7385b42bad463174532a3ef8fa6898e1e6654ded8e14a

Observation 541559a0-4b3b-41d3-a59b-ab3bed9e9af5 · outbound

This paper cites AudioLDM: Text-to- audio generation with latent diffusion models,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation AudioLDM: Text-to- audio generation with latent diffusion models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.957301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.509125Z digest=sha256:ab2d4509a45b30635c136055e149df51acdf6c7d1fe2a7a7dfbccc53b65b626a

Observation cd8b6482-9aa2-4ba4-bdb3-b51e71df9ff1 · outbound

This paper cites Denoising diffusion prob- abilistic models,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Denoising diffusion prob- abilistic models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.514122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.514122Z digest=sha256:9f9379598dce1ef9a66b409ce800b4de35a4b996590aa49de3f01c9510329797

Observation 9f7c8f5f-b1c0-46a9-8f02-bbc5532c1cf4 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.519045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.519045Z digest=sha256:08c10da4268b8c13cdfca345f1df0ba55483303d666ac51e89aeef6cf1ceb6a4

Observation 3c66992e-5c2a-4837-b6c7-8a726636d3fd · outbound

This paper cites Simple and controllable music generation,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Simple and controllable music generation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.524157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.524157Z digest=sha256:5656ce567e8cc4240fe8886f83aa8395e07a9ae9d0a509fff434ed0560bf3fc5

Observation 1027dca1-02c4-4226-af49-b809f9249866 · outbound

This paper cites Compa: Addressing the gap in compositional reasoning in audio-language models,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Compa: Addressing the gap in compositional reasoning in audio-language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.932511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.529642Z digest=sha256:6057dfd46b9b301907b1cb6af0edd249fe9a497b9fa5bda194945050d8da4064

Observation f45c1f0f-cdf5-4754-bc36-44144c16a673 · outbound

This paper cites A Demand-Driven Perspective on Generative Audio AI.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation A Demand-Driven Perspective on Generative Audio AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:48:46.737002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.534399Z digest=sha256:16a2c0e8b36e0e00674b1f80b2668e29f6326984d3f73048d55ec5f6dfcf9227

Observation eb686d45-1596-45c6-bdde-87fa5cdca707 · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.921005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.538457Z digest=sha256:1b8bc45c873241ec48068d878dbf9677bd9eae6ffa3dd65797172d246e1789a5

Observation 241993d4-bac1-4f8f-b8e1-547179c129fd · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.542856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.542856Z digest=sha256:ad4552a7d809911b69a605876b219021786bfb277a037902104c1a3f0e3c16e9

Observation ceaa1dff-4284-444d-a15a-1b22d0a23815 · outbound

This paper cites Generative adversarial nets,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Generative adversarial nets,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.909759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.547589Z digest=sha256:9e6f652b65ef7014974d9fa7aed883e4d7964686363ab10bd8d08b06c7500bd4

Observation b0e16031-dc20-45c9-8d62-7e7a487310bf · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Tacotron: Towards End-to-End Speech Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.551986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.551986Z digest=sha256:bc30e905df54b72332cdbd2bc8e732a55bbf514784143d6ab2511e2244dec2af

Observation ee3b2b77-d107-430e-95b4-d859b455c981 · outbound

This paper cites Auto-Encoding Variational Bayes.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Auto-Encoding Variational Bayes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.557063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.557063Z digest=sha256:c64385f8f4ad12deccceaad40c00cb4cf854a8a1d888d83ff8da16d877040419

Observation 0a677f8b-f682-4499-a7d2-9425e63bd858 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation AudioGen: Textually Guided Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.561780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.561780Z digest=sha256:0e860c1c1eff1d53c0df28eea04f2f3c28e75346e4f91c469145aca64f5604d6

Observation b1135e13-7d02-4d6c-aed9-4297887c124d · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.566564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.566564Z digest=sha256:fb292fd4a68596e311813bca2a0da31316c99d9da119569e5ba3dff7282ad2ac

Observation e0ff1037-5128-491d-b90c-23cad86a6c68 · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Diffwave: A versatile diffusion model for audio synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.898682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.570830Z digest=sha256:c896fca9626ee6a6ba5858fdb35a1a5f7eb86c9edc9deb135a3664628eebec10

Observation c68c40c9-a15c-4e56-bbfd-8a1fbd1f1769 · outbound

This paper cites Music controlnet: Multiple time-varying controls for music generation,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Music controlnet: Multiple time-varying controls for music generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.574553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.574553Z digest=sha256:9dde9394d9f3f8b4dbc0b5a9ce765417ce2c65c44772d6ef232cda1b27bb8ee3

Observation f9effb47-a044-4fbd-aea7-e2b9298c6750 · outbound

This paper cites Hierarchical Generative Modeling for Controllable Speech Synthesis.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Hierarchical Generative Modeling for Controllable Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.579511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.579511Z digest=sha256:e0b4671f34fbc9e7e9a5f4b4eb99ea88f35b8adea3e2872e352369b89450a1ba

Observation 9c4c451b-9dc2-4e8d-bfc3-967a60d57551 · outbound

This paper cites Deep voice: Real-time neural text-to-speech,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Deep voice: Real-time neural text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.878616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.584847Z digest=sha256:f1bcf44da7dcc1f0c3d96c2eb5105ccf4f5759ea51fbb3ecb7e02b95378d241e

Observation daa8b8fe-2100-419d-b792-eec6dbe7f438 · outbound

This paper cites GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.865812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.588821Z digest=sha256:834fa85dcdb89897c0759c7f2edc4e680cf9f6ca7b1ca0b751b976614e45b67e

Observation 1e336f2c-2986-4316-a703-4f19b1ece08f · outbound

This paper cites Mistral 7B.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.593367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.593367Z digest=sha256:b85d06b472e20a19927691bfdc0140944f2bca751cf50277f1c7cbde4c8f13c4

Observation 995a5bcf-71d9-437f-804a-6fd8080ad16f · outbound

This paper cites an unresolved cited work.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:48:46.852713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.598103Z digest=sha256:7adc74210c3a96111289ab409dafbb75b2672b03938c7aaf06a761d584dff6c3

Observation ee1ac531-ddcc-4e94-8068-b0df41594ee5 · outbound

This paper cites Crepe: A convolutional representation for pitch estimation,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Crepe: A convolutional representation for pitch estimation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.840699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.601685Z digest=sha256:8823d9970f6a771ecb6095cd17e1eb086c914a137f09f26467b3ff47c7b0be52

Observation 8a25975b-8b2d-401a-9183-ac18f509fab8 · outbound

This paper cites Suppression of acoustic noise in speech using spectral subtraction,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Suppression of acoustic noise in speech using spectral subtraction,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.829660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.605399Z digest=sha256:902c1175659f1401f10d748b8069f7f5fbfcbe50ecb74dc2e5117989a5be284e

Observation 27dcc94c-79c3-4d67-9e03-5f6f398fa878 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Fsd50k: an open dataset of human-labeled sound events,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.818296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.609249Z digest=sha256:c37130b99c0af70c159fa90f18c25994df82c720028395d0e6e703db64dfd1b3

Observation 10696c29-daa1-4fca-9a8a-aafbfc793e38 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt-enhanced diffusion mod- els,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Make- an-audio: Text-to-audio generation with prompt-enhanced diffusion mod- els,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.807508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.612562Z digest=sha256:89c76c02c723ae4722b6925131b439ba53563e89aee6c584bb9d01e2938574ac

Observation ffe3740a-9c03-453d-b962-963732da51d3 · outbound

This paper cites Scalable diffusion models with transformers,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Scalable diffusion models with transformers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.797395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.615943Z digest=sha256:70d457c09b3530cae2591f60d57c0f59de74d7f584996b42c5ea4e8b47263c27

Observation 672f74b3-78c4-459a-9ac3-164023fe7173 · outbound

This paper cites Scaling instruction-finetuned language models,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Scaling instruction-finetuned language models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.619020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.619020Z digest=sha256:c3a11f9aafbdc59b28f2c2915db48727a863719168be8ce9b388b61a357521a2

Observation ffd59caa-dbe8-4432-b86c-877f2b71d670 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation High-fidelity audio compression with improved rvqgan,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:48:46.779106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T16:48:46.622710Z digest=sha256:db455c79ee126a5e42265e2477cab4b03259d3cb1690209dce63cc56118a661d

Observation beb07327-682b-48f2-9c6b-b7122c355d1a · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:48:46.626775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:48:46.626775Z digest=sha256:1367672226ed84022313657c3de625344ce7ae68bbbe439b025d48fea144a308

Pith citing papers

Observation 3e4bfec1-d8dd-4d70-8845-cb42683db895 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.240224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:6049b8165be39819d40abd072d52013b37f9a8c20785f7c7f04ec47647391618