Pith. sign in

Paper Citation Record · LEDGER

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching

As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2608.12951.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12951 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:28.185416Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved54
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9672533c-2900-4e73-bf11-ffdb694aaf54 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Make-an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.737605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.737605Z digest=sha256:489e87a52bc613abdd34ade2f4063c07aa720a79c275bfd0220448bf1814857c

Observation a4b9f80f-36c2-4ce0-9b96-eab77a1aabbc · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.743079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.743079Z digest=sha256:9c1a23e61b6b48a7fec785248527d6a4964b452aeee77476ddef4b7ea4880d4a

Observation 207557e9-24ec-461c-bbb6-0a732429bc3f · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioGen: Textually Guided Audio Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.748642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.748642Z digest=sha256:c3a78e306a3cf92785e071166dd3176f30dd5c90f4cec91edc0aa72d3ad15fce

Observation 4387e8ab-111e-4a94-a289-b587784c521e · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.754377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.754377Z digest=sha256:a3a185829c2e31a8f68d7afa5372062d76911d66d0c44fa9ffc346fe10997867

Observation 85a6239e-927e-4b97-a397-a64ac48bcab1 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.763914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.763914Z digest=sha256:86806f9f47e45d6e609ca453cadb17680f46eb6c78e5da65ab9a793cb4dd3a9f

Observation d5d7220c-6f7b-4fc1-92d5-a9c9ae5d514e · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.785493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.785493Z digest=sha256:3467c3acbdc952355ce3d03794baad8b75d55ba81e3d882de681f1ad37d10a5b

Observation 93765a81-b0ca-4fe5-91bc-dddd88b5959a · outbound

This paper cites Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.831547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.831547Z digest=sha256:bbaba67e5b66ec657ad98338146307860f0ec78b21c71b3c9fe60ceceb9f74b5

Observation aec89b44-9c54-48c2-9f3d-cd637c4d0ea0 · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.851577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.851577Z digest=sha256:4b0ef4fc6c67a7ba74696a4bce4f32740ba14b024df4a1c2b388ccf9d832e007

Observation 6f4ee8d5-f4cd-458a-bdd3-7c872f719059 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.902369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.902369Z digest=sha256:3ad5e47d6e51bb6503ebc9e120c11e519d50f3c22790234a00161668a01dba6e

Observation 46d04179-0680-4767-b39e-7f675c76f47d · outbound

This paper cites Audioldm 2: Learning holistic audio generation with self-supervised pretraining,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audioldm 2: Learning holistic audio generation with self-supervised pretraining,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.908184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.908184Z digest=sha256:f9d9c69c7d63cf1b84536203057c09520ff9ba1b9f78e6c6c19c93f78619ca4d

Observation 0098b65d-aebd-41f6-95de-d804a312b0de · outbound

This paper cites Taming Data and Transformers for Audio Generation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Taming Data and Transformers for Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.912849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.912849Z digest=sha256:8458050468e0d9df5c99660c3d8ec8c1ea7f3ea71a5aa4b802b553197d38e1d4

Observation 703f43e4-0dd6-47ff-ac98-e9c8560680f6 · outbound

This paper cites EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.917909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.917909Z digest=sha256:395103a1643305d60f7c6ec28876921dba64211aef460e1c5b4a61992b901a72

Observation 832cc07c-bdf3-4b29-b0a0-36c2bb73b611 · outbound

This paper cites Stable audio open,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Stable audio open,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.663831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:27.922810Z digest=sha256:b4312c45ca0bf3666402947f985f9a4e80fdf4eceb3faad552b09fb16d021013

Observation 4747fb0d-534b-4f98-9c60-4397cd3c608f · outbound

This paper cites Mmaudio: Taming multimodal joint training for high- quality video-to-audio synthesis,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Mmaudio: Taming multimodal joint training for high- quality video-to-audio synthesis,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.927761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.927761Z digest=sha256:5231fcf2a11afc0b861af789ad2cb3083a95ce8bac73c0aa6b23dd0c8cb15b12

Observation 45671c5d-729c-4b2f-9311-fd5655d35707 · outbound

This paper cites AudioX: A Unified Framework for Anything-to-Audio Generation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching AudioX: A Unified Framework for Anything-to-Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.932491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.932491Z digest=sha256:5b04720e4b27f6b3e9c108e7c1aacadc7ab4d1d06fb0d3b459403c27c3b1ab5d

Observation a25dad6e-f566-4d06-bebc-ffb4c48e1d72 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.937455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.937455Z digest=sha256:a000dc78c4bbda1dc663cff87a2a2668b62636465e2876f1d20122af92f7b41f

Observation 5cb84e5d-54e2-42f3-92a6-71ffd13150fd · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.943160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.943160Z digest=sha256:6e5dadf7eeb68bdc2c32bf72ad350b90102ce13225f9aaf973f4cd14607f15c1

Observation 2beaf4cd-24d5-4291-bbdc-1d2f90e24be1 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.948142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.948142Z digest=sha256:5e3a73eb464207402c55ca00415bd1b9cacd71f6bb68a9f7d15a87ae30f9fa7f

Observation 62f664d0-1168-4fc1-a8ad-4db1a983aec4 · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.953209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.953209Z digest=sha256:1c4dd1d4e49807ebdb48cbdb5d0558aafd282dba259479f22ad3d5ffd4dcf58e

Observation dfd7a58c-6806-4fd3-ad39-28cbe7446353 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.958086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.958086Z digest=sha256:592a48ca664cadd2827a3442ed57b817939009198a4656b3e20715be7c6346b0

Observation d034211d-1497-4203-8935-b689d20a470c · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.963256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.963256Z digest=sha256:0e5210763a127be358b811de676104d5d5c47471863cd53b29c7d4a4ea0f2be8

Observation f0c5aa2d-014c-4794-ada7-72b7a84465f9 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.968462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.968462Z digest=sha256:22d2c4c8e4c24eb18a28a440edb65cf40cb2358a258ebbb3def0b6f990e8e0c7

Observation db15bfb1-cc3a-4853-9f7b-47dc385ff421 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Ditar: Diffusion transformer autoregressive modeling for speech generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.973803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.973803Z digest=sha256:41d857687a71bcdde22b5b99347c287113fe3630c2fd0a877c6d87217c65bfbb

Observation a5b01b59-4c4b-4f1a-a1f9-9e80d93ff61e · outbound

This paper cites VibeVoice Technical Report.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching VibeVoice Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.980409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.980409Z digest=sha256:4f1c3114300a4981f345e4bc95e6133cfc4127fb7fc57ebaab365d1ceec08f82

Observation 8bb279a8-fcfe-4d41-9a00-afdaa650f374 · outbound

This paper cites Autoregressive image generation without vector quantization,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Autoregressive image generation without vector quantization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.608699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:27.985469Z digest=sha256:d96920af12b90876109714ae8869d52ebaf79486d4cfddc1efaac44a33c6de10

Observation b197548f-8178-47c0-8266-9e71eb52b8d8 · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.991313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.991313Z digest=sha256:10291c8d9b18a13ebee17022fad862b1f4202c4ebed85738cfa281f46537ebe7

Observation d3566ea0-abd0-408d-b501-4b91df69ad02 · outbound

This paper cites Diffusion forcing: Next-token prediction meets full- sequence diffusion,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffusion forcing: Next-token prediction meets full- sequence diffusion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.588953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:27.996597Z digest=sha256:e507c4a4dad23a0b3a4c505999d21d74cb7823275c081a5034b9e3984e9a3048

Observation 75d5a000-c748-4015-bb05-8ee6f530eebb · outbound

This paper cites Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Ar-diffusion: Asynchronous video generation with auto-regressive diffusion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.571951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.001526Z digest=sha256:b626664fd5e7720e7a808b5faedb1ec4e1b77f0814b36711bc461f17ab5cd570

Observation af3eded2-31b4-4c50-837e-dfed01a0f30d · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching From slow bidirectional to fast autoregressive video diffusion models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.006335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.006335Z digest=sha256:3f5182c40f5f4ca5fb6fcf6d18d6e5d620488502a92bd4368d91c874346b4446

Observation 07a195b0-f4f5-4445-9bed-70a0686ac9d2 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.010553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.010553Z digest=sha256:d8ad0905408575f171d2f9271fbfbe1f27275141d91ed0086cacf85906edbf19

Observation 1f3e1cfa-ff5f-4b5f-8018-3e59180656f4 · outbound

This paper cites Training language models to follow instructions with human feedback,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Training language models to follow instructions with human feedback,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.014885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.014885Z digest=sha256:52236fb87056e5b44222398fb9c114c22823864d2bf0be927da91e6fdf55a784

Observation f2d2b4e9-7be7-4377-84d0-03be0240496b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Direct preference optimization: Your language model is secretly a reward model,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.019584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.019584Z digest=sha256:e465a24fd82e3c4f877dde9164b5d40e356b503ea476b26580fe3d3d9a4a44da

Observation 7e62f14f-8aa5-47c7-b3a0-665605ff195b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.023555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.023555Z digest=sha256:e94bdbd2ba56241bcd6e9ba5b72da083275dcbcacc629ba0e4b2349c214729e8

Observation 9ba2ffbe-557a-48f1-9681-257523f30d75 · outbound

This paper cites Diffusion model alignment using direct preference optimization,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Diffusion model alignment using direct preference optimization,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.028855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.028855Z digest=sha256:643c9d66c35288ecb93c0a32f2c4b95a9983921dd161cd6f44b11202f3ece4bb

Observation c80060bc-4f9c-4e80-aa63-5d53c0bc9620 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Flow-GRPO: Training Flow Matching Models via Online RL

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.033504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.033504Z digest=sha256:58d8b18d3d6c36623981154b5e8a3ca84b7deaf6ee4fb2b78d1a7e8bb815df42

Observation 66b75809-d0c6-4cef-9110-77544283a6d2 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.038349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.038349Z digest=sha256:d7b6d10987c5e052b2b77ac98c5146bd3b66999f7a888cf19b5b8620affcecb9

Observation 4105aa1f-213f-48c2-8a20-1390f1db9ce3 · outbound

This paper cites TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.042751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.042751Z digest=sha256:573130cf526672b7a0edf68a730dda085a54175c01d755bfc8b7184af07f7363

Observation 82799c0c-fb11-4167-810a-2fa0dba1358e · outbound

This paper cites Prismaudio: Decomposed chain-of-thoughts and multi- dimensional rewards for video-to-audio generation,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Prismaudio: Decomposed chain-of-thoughts and multi- dimensional rewards for video-to-audio generation,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.047622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.047622Z digest=sha256:974c6046084d8b2e5573b246edbc5473b1c1ddb2009e0fa7c7da49de300ca6eb

Observation 51032a5c-5bd4-4272-bfe1-26817814fcce · outbound

This paper cites Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.052309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.052309Z digest=sha256:cd8c772013ff11d5f757ce18c945ab5cc20d92297ff46e6cb04a52ad5b0ee9cc

Observation b5083e75-2e93-4974-ba49-3b99ad272108 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Scaling rectified flow transformers for high-resolution image synthesis,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.057140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.057140Z digest=sha256:ad8737bc8c807c82b40f580ccea3adc2f923618deb92ec025b2ff8f80485a63e

Observation 75e931a7-5bd3-4985-b1dd-2a0d92bbb7e8 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.062233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.062233Z digest=sha256:7bed4eee65775ef0ef3ad1f7f07b2cc57862a03a31fbc7578416faf7852f870f

Observation 41fca24d-5715-4638-b59f-0c3a15bdffc9 · outbound

This paper cites Pushing the frontier of audiovisual perception with large- scale multimodal correspondence learning,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Pushing the frontier of audiovisual perception with large- scale multimodal correspondence learning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.067072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.067072Z digest=sha256:988c17b531a186a24e176c9fa057914826aaa29099750f4e4f5d0b74b4727a70

Observation 6cc44144-8abd-4bc0-8e6e-cac480c7ee4b · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Robust speech recognition via large-scale weak supervi- sion,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.071833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.071833Z digest=sha256:e2249beeffdfbdbfb0e97a5a4c40802a351b4cb61db65246935953d2c435548b

Observation c2679713-1b2e-4c87-a486-2d476399e67d · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.076577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.076577Z digest=sha256:f225282bc580b6f14da7afcc467bf7f32366a1791226ff4887835162c4a7495b

Observation a5cf7cb9-6724-4db5-9aae-ebf7f988b05a · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Perception Encoder: The best visual embeddings are not at the output of the network

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.082042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.082042Z digest=sha256:613ea706b8e8989c926fd16c63da7e46226bb8c205624419318d550ab3ca3a35

Observation 57793283-bf3c-49a9-99ec-c25775040d89 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiocaps: Generating captions for audios in the wild,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.087451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.087451Z digest=sha256:c5182be1749d5be78d3ea71a40516e5da6f9c209011d023ae2dd3a0d0799c553

Observation fc60f91d-1943-4301-b70d-ba330f6667fc · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.093470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.093470Z digest=sha256:14deae9c7e30b35bed3945a5611371cdf7c8e7c8976364a7a503f23ce2dbe44e

Observation 7d63786f-1a21-4db8-8bdc-9979a0468d25 · outbound

This paper cites Qwen3-Omni Technical Report.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Qwen3-Omni Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.098401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.098401Z digest=sha256:c5d5fd393a7cb4c7f56c58b91a349ce9c8c68d66ce23a3c27923c7be90049742

Observation ed244f27-16a0-48bc-b0fd-649b05ff085b · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.103563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.103563Z digest=sha256:930956c92eaf35099c252eafca8d31a647afd45c47f48dc1e811076a5cc3feae

Observation 1ce6a461-38e5-4e33-923e-615220c5ba34 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.108913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.108913Z digest=sha256:b4eec5cd0b32df1565d9fdd85c86182f67b35f4b21e1085845e2f2d0a5d10e7e

Observation 4d226849-db7e-46e5-b47a-528492a59207 · outbound

This paper cites Stable Audio 3.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Stable Audio 3

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.113594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.113594Z digest=sha256:b3483229b7cceb86c30511adf44097b1fffe9ad0c8a68431ac88d6fe3193b21d

Observation 80503fad-106d-4373-b186-6af0fc741e39 · outbound

This paper cites Qwen3-TTS Technical Report.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Qwen3-TTS Technical Report

Reference 53

Resolution
malformed identifier
no resolver link, observed 2026-08-15T19:59:28.118525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.118525Z digest=sha256:e43be2960672c4b3ef0b6929dd7e2874686c42155395f31fa4eb2fa4cc927833

Observation 03b9926f-c8a4-435b-b776-bcb75964a85c · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audio set: An ontology and human- labeled dataset for audio events,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:28.123447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:28.123447Z digest=sha256:9ac072361dd42c8ec73615d332dcf69c21914d498e405bad01cf423a389518fe

Observation 71fd00e1-91d5-4e27-9f7b-fd3f40df3c7d · outbound

This paper cites Vggsound: A large- scale audio-visual dataset,.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Vggsound: A large- scale audio-visual dataset,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.411187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.127970Z digest=sha256:337ed13f43b60e994f7d01210bf757a3cdbd3b00f14778d1851ed874e46c220f

Observation 87724f78-a4f8-4606-8c90-d3110b0ca09c · outbound

This paper cites Do not imagine or add details not mentioned in the source.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Do not imagine or add details not mentioned in the source

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.387845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.132670Z digest=sha256:4e1b6d628b5357e0ef5c160e920662d7deb40f16b12674f6c6a774a851625928

Observation 816fc41c-7a99-4bfe-8ed7-08ddcfb311ce · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.352341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.136808Z digest=sha256:75491daa71a4d10dd84744dc7c40257707e283f00d399ad0dc1385ad086d9d63

Observation f888248b-5156-4dfd-b180-01c5be2dbef2 · outbound

This paper cites Preserve timing and dynamic changes.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Preserve timing and dynamic changes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.328448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.140856Z digest=sha256:fe5f1e850919ae654f63260bb684cee49c7640572c17eedbb2ba94e8aa3c6b4c

Observation 99fce691-8cae-4cdb-a274-449959ed1b9a · outbound

This paper cites [SPEECH CONTENT](Transcription).

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching [SPEECH CONTENT](Transcription)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.300221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.145387Z digest=sha256:3b095db077bd759507501d2ab1994cff7d3e1c7cbeeec1ee94af5c583e009f7c

Observation 8b49d9ef-02ad-4c33-ac26-5f4555586862 · outbound

This paper cites caption short.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching caption short

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:59:29.281249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.149459Z digest=sha256:2b5d397e564343d2c853bb94c6b69d08b0a4f79854fb16005112e9bc0f47e9a7

Observation 5a1aaba8-c8d6-4254-8d74-3f72e442175a · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.263787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.153835Z digest=sha256:6fd8901f07438584d87ca45bbad191c8bd5f6a70327b5b0e80271aaac98f7e56

Observation 72c075a1-0cac-41ff-ac84-5e4a81757ab0 · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.246213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.158610Z digest=sha256:49e1bf9b4c44691927b30b98241fb2fdec53732e2eb6ba1b59cee4489afa7a4d

Observation e3d46a85-a6e5-4a5f-9391-385e68dff3be · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.230376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.163731Z digest=sha256:65dd05357673f0fd33c31242c93961c3143bd4bb96eb9b1f1a5f3d6ad2f24389

Observation ec19f2be-e176-4b89-b258-0136fb59b782 · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.214950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.168888Z digest=sha256:aee4bffab36b18b53c4e3374d801bd96c847202c40b56c498536c39ba3b088e5

Observation 35bd4d4f-b5c6-4b43-b5ab-7a89b4847526 · outbound

This paper cites an unresolved cited work.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:59:29.197938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.173507Z digest=sha256:c82577114974dfed15b91f214aa169ab90872be8f6adfcf185ec5f13a560cac2

Observation 064dc0c6-3cb4-4d83-8434-2fce47717b4b · outbound

This paper cites candidates.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching candidates

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:29.180424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.179211Z digest=sha256:a237bd2cd969e675c5eaf0c7c6bab52a65fba6afc4baf27b4775e678cb4c9a81

Observation e87d1c4f-25ce-4ce7-84d3-5ff3e35979f2 · outbound

This paper cites TABLE VII PER-CATEGORY RESULTS ONMECAT-EN.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching TABLE VII PER-CATEGORY RESULTS ONMECAT-EN

Reference 67

Resolution
verified exact
raw_fallback, observed 2026-08-15T19:59:28.323919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:59:28.185416Z digest=sha256:baa2186e9893eea53cd393389ccf3c0e9e047e6d083cb2effc1cec868eb58458

Pith citing papers

No inbound Pith citation observations are available.