Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:28.185416Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.12951.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:28.185416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9672533c-2900-4e73-bf11-ffdb694aaf54 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b9f80f-36c2-4ce0-9b96-eab77a1aabbc · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 207557e9-24ec-461c-bbb6-0a732429bc3f · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioGen: Textually Guided Audio Generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4387e8ab-111e-4a94-a289-b587784c521e · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffsound: Discrete diffusion model for text-to-sound generation,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a6239e-927e-4b97-a397-a64ac48bcab1 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d7220c-6f7b-4fc1-92d5-a9c9ae5d514e · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93765a81-b0ca-4fe5-91bc-dddd88b5959a · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec89b44-9c54-48c2-9f3d-cd637c4d0ea0 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f4ee8d5-f4cd-458a-bdd3-7c872f719059 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46d04179-0680-4767-b39e-7f675c76f47d · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audioldm 2: Learning holistic audio generation with self-supervised pretraining,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0098b65d-aebd-41f6-95de-d804a312b0de · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Taming Data and Transformers for Audio Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703f43e4-0dd6-47ff-ac98-e9c8560680f6 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832cc07c-bdf3-4b29-b0a0-36c2bb73b611 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Stable audio open,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4747fb0d-534b-4f98-9c60-4397cd3c608f · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Mmaudio: Taming multimodal joint training for high- quality video-to-audio synthesis,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45671c5d-729c-4b2f-9311-fd5655d35707 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioX: A Unified Framework for Anything-to-Audio Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a25dad6e-f566-4d06-bebc-ffb4c48e1d72 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb84e5d-54e2-42f3-92a6-71ffd13150fd · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2beaf4cd-24d5-4291-bbdc-1d2f90e24be1 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f664d0-1168-4fc1-a8ad-4db1a983aec4 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd7a58c-6806-4fd3-ad39-28cbe7446353 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d034211d-1497-4203-8935-b689d20a470c · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c5aa2d-014c-4794-ada7-72b7a84465f9 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db15bfb1-cc3a-4853-9f7b-47dc385ff421 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Ditar: Diffusion transformer autoregressive modeling for speech generation,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b01b59-4c4b-4f1a-a1f9-9e80d93ff61e · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching VibeVoice Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bb279a8-fcfe-4d41-9a00-afdaa650f374 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Autoregressive image generation without vector quantization,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b197548f-8178-47c0-8266-9e71eb52b8d8 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Multimodal Latent Language Modeling with Next-Token Diffusion
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3566ea0-abd0-408d-b501-4b91df69ad02 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffusion forcing: Next-token prediction meets full- sequence diffusion,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75d5a000-c748-4015-bb05-8ee6f530eebb · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation af3eded2-31b4-4c50-837e-dfed01a0f30d · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching From slow bidirectional to fast autoregressive video diffusion models,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a195b0-f4f5-4445-9bed-70a0686ac9d2 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3e1cfa-ff5f-4b5f-8018-3e59180656f4 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Training language models to follow instructions with human feedback,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d2b4e9-7be7-4377-84d0-03be0240496b · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Direct preference optimization: Your language model is secretly a reward model,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e62f14f-8aa5-47c7-b3a0-665605ff195b · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba2ffbe-557a-48f1-9681-257523f30d75 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffusion model alignment using direct preference optimization,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80060bc-4f9c-4e80-aa63-5d53c0bc9620 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Flow-GRPO: Training Flow Matching Models via Online RL
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66b75809-d0c6-4cef-9110-77544283a6d2 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4105aa1f-213f-48c2-8a20-1390f1db9ce3 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82799c0c-fb11-4167-810a-2fa0dba1358e · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Prismaudio: Decomposed chain-of-thoughts and multi- dimensional rewards for video-to-audio generation,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51032a5c-5bd4-4272-bfe1-26817814fcce · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5083e75-2e93-4974-ba49-3b99ad272108 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Scaling rectified flow transformers for high-resolution image synthesis,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e931a7-5bd3-4985-b1dd-2a0d92bbb7e8 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41fca24d-5715-4638-b59f-0c3a15bdffc9 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Pushing the frontier of audiovisual perception with large- scale multimodal correspondence learning,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc44144-8abd-4bc0-8e6e-cac480c7ee4b · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Robust speech recognition via large-scale weak supervi- sion,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2679713-1b2e-4c87-a486-2d476399e67d · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5cf7cb9-6724-4db5-9aae-ebf7f988b05a · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Perception Encoder: The best visual embeddings are not at the output of the network
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57793283-bf3c-49a9-99ec-c25775040d89 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiocaps: Generating captions for audios in the wild,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc60f91d-1943-4301-b70d-ba330f6667fc · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d63786f-1a21-4db8-8bdc-9979a0468d25 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Qwen3-Omni Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed244f27-16a0-48bc-b0fd-649b05ff085b · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce6a461-38e5-4e33-923e-615220c5ba34 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d226849-db7e-46e5-b47a-528492a59207 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Stable Audio 3
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80503fad-106d-4373-b186-6af0fc741e39 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Qwen3-TTS Technical Report
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03b9926f-c8a4-435b-b776-bcb75964a85c · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audio set: An ontology and human- labeled dataset for audio events,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fd00e1-91d5-4e27-9f7b-fd3f40df3c7d · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Vggsound: A large- scale audio-visual dataset,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87724f78-a4f8-4606-8c90-d3110b0ca09c · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Do not imagine or add details not mentioned in the source
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 816fc41c-7a99-4bfe-8ed7-08ddcfb311ce · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f888248b-5156-4dfd-b180-01c5be2dbef2 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Preserve timing and dynamic changes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 99fce691-8cae-4cdb-a274-449959ed1b9a · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching [SPEECH CONTENT](Transcription)
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b49d9ef-02ad-4c33-ac26-5f4555586862 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching caption short
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5a1aaba8-c8d6-4254-8d74-3f72e442175a · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72c075a1-0cac-41ff-ac84-5e4a81757ab0 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e3d46a85-a6e5-4a5f-9391-385e68dff3be · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec19f2be-e176-4b89-b258-0136fb59b782 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35bd4d4f-b5c6-4b43-b5ab-7a89b4847526 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 064dc0c6-3cb4-4d83-8434-2fce47717b4b · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching candidates
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e87d1c4f-25ce-4ce7-84d3-5ff3e35979f2 · outbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching TABLE VII PER-CATEGORY RESULTS ONMECAT-EN
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.