Pith. sign in

Paper Citation Record · LEDGER

Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2304.13731.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.13731 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:13:15.771340Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

15
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 69860cd6-e039-4abb-818e-5f10f765f5be · inbound

Generative Semantic Communication: Diffusion Models Beyond Bit Recovery cites this paper.

Generative Semantic Communication: Diffusion Models Beyond Bit Recovery Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:44:13.902413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T08:43:17.500704Z digest=sha256:6300e32871271178496fdc543fa9014f7816717450e5a37865225671f0a7924b

Observation 1c1ce906-c66e-457a-9914-2eec6b7bf2f4 · inbound

Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling cites this paper.

Training-Free Multi-User Generative Semantic Communications via Null-Space Diffusion Sampling Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-24T01:43:43.467983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T01:38:42.441433Z digest=sha256:105c13282cc74051b1785928b10e86d936e7192a6b5ee51cc8d1d3a6ab97146c

Observation 707a0d2c-634c-4acf-a713-b365b0ec29d6 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.275961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:e7a6752b71c50f4ead1ebf4946aba6a4fa2135edc6407f5c5f7b62a7b92ed0f1

Observation ca53461d-b78d-4ea0-8b5e-812a18c0efc1 · inbound

MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI cites this paper.

MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:27:20.859159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:27:20.859159Z digest=sha256:159df3d317551850d8c6293b059cff22c1a320f7d4fd19b4d9cd9710f54d7979

Observation 060332f0-dc0d-4512-9549-4f8f5ab1e690 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.398258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.398258Z digest=sha256:1940966351b57b1ef84c66f6a4f3c2bdc49601c5766a1503d784c513acf7210d

Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · inbound

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation cites this paper.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.881954Z digest=sha256:cb45f6fc43f7bdc5af0428e5b8a8cae2122e0a44c535da2e56399363f780510c

Observation 9ea5d44d-f34d-4927-b625-ef939d97f6f8 · inbound

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis cites this paper.

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:35:57.679504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:35:57.679504Z digest=sha256:6643a8757773db88055fbdc9b28d7cc1fe247a0689e0dee1a59eae2b4c36d2b4

Observation fee7105a-96f9-46eb-bcba-338e0d6ac279 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.520734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.520734Z digest=sha256:9bcc3901955ce9506e627fb494eaafcbff84fd26273737247def0d536cfcff4c

Observation fa2e99e3-c5b0-481b-8d65-5d64ab32c365 · inbound

Neural Vocoders as Speech Enhancers cites this paper.

Neural Vocoders as Speech Enhancers Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:59:21.026366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:59:21.026366Z digest=sha256:e1f66f64718ad42eeb7f402e0546850f10697c412baba82025cf0e38af4e3b7c

Observation f7d37e2e-d5cb-4911-88fd-5825de27fe83 · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.752505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.752505Z digest=sha256:96d65b54ef918352567eb4d5dcb9cf200e4847a1a42439921dfd11e8e4cc1279

Observation 652e3360-2970-470c-80d9-3d11db5d2752 · inbound

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey cites this paper.

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 285

Resolution
unresolved
no resolver link, observed 2026-08-10T04:36:38.383890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:36:38.383890Z digest=sha256:9745c967b156e7894618fbe9335c156c01ae44d4c495b7b276211a1357891a66

Observation cb759659-6ea3-46ce-a086-8709c5d47140 · inbound

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification cites this paper.

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:21:17.307566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:21:17.307566Z digest=sha256:68ad6dc0eaa4150b26fad661f317116cfe5f894f1b191edfec5757ca5c04d6f4

Observation ca845220-fa6c-49bd-a730-253acc8ed52f · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.239173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.239173Z digest=sha256:726cb0ecca5fbb3f25f6bcde7e6fdb4ce606e9a4e00d05994c0983c1752e1732

Observation 537e01b1-1e2e-4d28-8e31-4aaebdd6d25d · inbound

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation cites this paper.

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T20:12:29.425344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:12:29.425344Z digest=sha256:938b8cc2f3b3de74141e775ff0001b75c037fde2f7729b86c7c364b5ef824ff9

Observation 3ac16a45-5bd5-4f75-a9ae-0834e90d4fc4 · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.502166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.502166Z digest=sha256:eec81468edd0cb2c90c85c03ea0b29cee44a43ebafa97d40b5a65a07a9a9a020

Observation 67958216-8203-4b5c-bd0c-bd450947469c · inbound

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding cites this paper.

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T20:14:03.585698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:14:03.585698Z digest=sha256:4d5ee3e8748c4187951776c01f19497d82007f2003361b716f9fd72b6e69ab14

Observation f073097f-4833-44c7-868f-e8d33fb37fc5 · inbound

Rethinking Score Distilling Sampling for 3D Editing and Generation cites this paper.

Rethinking Score Distilling Sampling for 3D Editing and Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:13:15.771340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:13:15.771340Z digest=sha256:de9f4b2c23f11ecd7bbbb59e3a3fdb503a6c0de5c9a92c5c4799f268ee531048

Observation 6f60b297-9e23-49ca-8932-94b1bd89c1c8 · inbound

NeuSEditor: From Multi-View Images to Text-Guided Neural Surface Edits cites this paper.

NeuSEditor: From Multi-View Images to Text-Guided Neural Surface Edits Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:58.297358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:05:58.297358Z digest=sha256:76f037688d8f81c8e81f98a8188a9f0636e4e77fa68abcaad85c8baea5d66747

Observation 039e2613-8da5-4e80-a2dd-bbb982f6bfe7 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:24.909538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:24.909538Z digest=sha256:9f59a59ae5657a2398b772a696ef41144a62c8eec27b602724c132868e82c476

Observation 8cc0151b-43c8-4f68-a72c-50d218cdacc8 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.192473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.192473Z digest=sha256:072edcab1604f9e5b21c8d2548eca7aba3fc84c5e3321f4e3979bcb2cb31a017

Observation 90dc13ef-de70-403c-83d4-b6dcc99ece2f · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.504432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.504432Z digest=sha256:8d9eff8e8a924ac3b6cb46bce9fdadee61e48b36a02fad2b4a0fbc69d6cc7192

Observation 77d4c176-df34-4284-a3c6-df2bc1bf91ef · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.386327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.386327Z digest=sha256:80a0bad18a1f982d20dbcbea4701d256d898d88cced541932f4c68db2dbcaea3

Observation c73a878b-9120-4d7c-b9d8-0d069b9b32f4 · inbound

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models cites this paper.

TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:37.929378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:37.929378Z digest=sha256:272ba969e16aef65c1dbdd8fca58593f404129faa31980cf3925151ad5dcd09c

Observation eb26b8d8-ce1c-4a41-a413-4f98fe9b0774 · inbound

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models cites this paper.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.694499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:463da17c927e20170a144a4595caefe85a542fe823a6ec9960923839bf2eaa76

Observation 8205c891-3462-4416-a8a8-d7edb75a6565 · inbound

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance cites this paper.

AudioMoG: Guiding Audio Generation with Mixture-of-Guidance Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.841838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T13:10:18.700497Z digest=sha256:e2511231d53f693979e8f63813591f98f6a378163e1884e139413f5855df125d

Observation 6dc1f721-c254-40e8-bbb1-550d02063719 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:31.996283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:31.996283Z digest=sha256:5bb843e9462212fec96f3b1eddac90d603a057e2624a53d306fae4f2af6dfccf

Observation cbe1b4c2-5e26-453d-bb72-c04d068231c1 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.035503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:8415c007b56c61f130cf487430cbc9b4157720b57bb0c0497433e2d31a0ff6d7

Observation cc2f9143-e0da-45f8-90c0-5b8af801aee5 · inbound

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts cites this paper.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.497266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:297ea8e269497587c28bc9abb58c75a172e34056b05d96cc5cab04e02642560a

Observation 08905fd2-4446-4e2b-ac4a-c5d44b33ca3a · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.943293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:0e4f27f55b4ce1a97018d046bd332478ae5b749ae6b95be85e6ed76ba5be7087

Observation 8b57c942-0463-4f7b-86f1-960219ace007 · inbound

Auditing Training Data in Generative Music Models via Black-Box Membership Inference cites this paper.

Auditing Training Data in Generative Music Models via Black-Box Membership Inference Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:16.272877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T08:45:23.523801Z digest=sha256:3752fc33d88188cea434a9e271b8c9943c2d5b2997fbc2b6f782d0971adea952

Observation 3fc835a8-7169-4cb2-8a84-a653542dc1b0 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.725940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:a3ab7ac7e20af16e6bc5779cfe43c83d568d7ac924ce3c2ebc03cd8e10977894

Observation cd8dac01-cc08-4c03-b8de-40e5cddd78c1 · inbound

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models cites this paper.

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:19:57.089048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:19:57.089048Z digest=sha256:4616fb42430de35e0e439c20f64f170914c68f750f5e3bab605c332140d10c1a

Observation a07c5b1f-b16c-444f-af20-39dc2fe7e313 · inbound

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration cites this paper.

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T17:47:15.593693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:47:15.593693Z digest=sha256:ec8cc091f045256a469450b5fe50c65f44ae347cb455e1ef2be3319da5560760

Observation 105e8c7f-b948-40ec-9d5a-68954b12baaa · inbound

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model cites this paper.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.896826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.896826Z digest=sha256:e54178604dc89e44932447cab39949ba7d2b62282baadcb13f943189c2b31ff1

Observation 1ce6a461-38e5-4e33-923e-615220c5ba34 · inbound

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching cites this paper.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.108913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.108913Z digest=sha256:837ba67226a5fe8e97dbf7f338fc3ba90f9089e891c9428d76163babc7b41af1