Pith. sign in

Paper Citation Record · LEDGER

Improving Text-To-Audio Models with Synthetic Captions

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.15487.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.15487 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:45:19.794354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.458768Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b58e351c-68d9-4742-a8bd-d02e5e6ce317 · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Improving Text-To-Audio Models with Synthetic Captions

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.794354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.794354Z digest=sha256:8d424301c1386269d22a53510c74b703486922853b8920a44f5c890af30bcba7

Observation f156cb9d-4ed3-4cd2-85d6-228143149360 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Improving Text-To-Audio Models with Synthetic Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.578346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.578346Z digest=sha256:5cd19cdd4f4b09e0209f0f65542d197283beaffeeeb213df7dac1173f63dae39

Observation 6257149f-256c-4648-b54d-2de605277656 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge Improving Text-To-Audio Models with Synthetic Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.738918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.738918Z digest=sha256:520b3b83735d62cae8c4148c08352b8d6b8bbd80082a2bdaa73d4ee3f809ce86

Observation 98d0d72c-814d-448f-96be-2c18c456fb85 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance Improving Text-To-Audio Models with Synthetic Captions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.857545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:cf6cd7e76fd0abd6e2c42e60b8c2b2bd2d707029d9d5c84d63a380694118f992

Observation 054a7420-fe6d-4f27-8300-7759615d491a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Improving Text-To-Audio Models with Synthetic Captions

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.460041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:f2334a62fc23b339abe91fb2a04d7e6a33f11accb2dc6f5f3cb1378419dc0b50

Observation ef1d5c12-b559-403f-8536-2acbeaddb5d6 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Improving Text-To-Audio Models with Synthetic Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:68d435eb221efd1e385913ea4c0639314b0fcd899a26b40245ebda8384b72824