Pith. sign in

Paper Citation Record · LEDGER

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

As of 16 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 10 inbound Pith citation observations for arXiv:2412.15220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15220 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:27.453564Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:07:18.202024Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.931162Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d715897c-cb03-4450-84b5-333207ec8dfa · outbound

This paper cites MusicLM: Generating Music From Text.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.368260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.368260Z digest=sha256:cf9773be68cde5ea6063833930ea7c2608af8043629597347862ef061b814a54

Observation abb54f07-cb9b-4d0f-aa35-d6e128898036 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Imagen Video: High Definition Video Generation with Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.393605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.393605Z digest=sha256:977dc3a4dcf0cf3455a2b52155065a7fc36aaa89c1671a7c30652d6494878698

Observation f5540ced-a70e-4c15-89a0-b2b79ce3f1d9 · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.398861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.398861Z digest=sha256:26e71262c68321692d60bddd952e8c42181b1656789a8000ce4eca19a4c76e70

Observation a294bb5e-6dbd-4118-9daa-0d89e0471be6 · outbound

This paper cites FoleyGen: Visually-Guided Audio Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text FoleyGen: Visually-Guided Audio Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.408048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.408048Z digest=sha256:771eba3549638b69dc4dc17575fe9366c4b29585fe8cd2c8881ecd16d6e5d443

Observation 0d3da6c7-1d27-4941-8bd5-a2b01f48b431 · outbound

This paper cites Text-to-Audio Generation Synchronized with Videos.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Text-to-Audio Generation Synchronized with Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.412504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.412504Z digest=sha256:e8d359f2374ac9e3a0e6e7a73267c5c85d8e0c9bf11aa84bb3427e42f62f1cc5

Observation e45ea5c1-09fa-4992-a19d-e2928f636e11 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.421876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.421876Z digest=sha256:1f90d189b3126159952181212fa8e22ac1534b1c57973c21172a8e0c9e2682af

Observation ca96df05-847d-4ac9-b96a-5c33e1c50ce4 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.426158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.426158Z digest=sha256:a3e0eaf4b24eea12bb0fcc784df768e4f2879e8121a26aeab9a77bde92e24f7a

Observation 513b7a70-aa2e-4386-9097-a4c11b164d9a · outbound

This paper cites NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.430832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.430832Z digest=sha256:06a0b40bb962d0c1ca973c6546cad7360f904880acb32680d620b8eaa027f701

Observation 187f45ca-e2c5-4f3d-ad1f-92ef1785f6a7 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Diffusion Models Are Real-Time Game Engines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.435257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.435257Z digest=sha256:cb6b2b3e2e9968e61733bc42f603ff72fc3ab06b380a17c0633b628fd8415a3b

Observation 16951b98-ed79-4e0b-884f-4bed32300366 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.439825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.439825Z digest=sha256:685c4d620c04f805180b10c09542b4f40191c8fe8ca02251e97ea0a753786f77

Observation 1398e0fc-7d6e-480d-9766-a9ac6e4bbea7 · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.444425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.444425Z digest=sha256:a3c6b58c28d51a917443be70299d911bf0636b96db13b143a7c94b743e22f856

Observation 1b4b0a64-75a7-4fc7-b916-7aa38610716d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.448864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.448864Z digest=sha256:798ac8b9744f4201167cb97312864c07502302259041485151fc3fa8f30cb74a

Observation de8052fc-ceeb-4a62-a3fc-39411cc372b1 · outbound

This paper cites FlowSep: Language-Queried Sound Separation with Rectified Flow Matching.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.453564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.453564Z digest=sha256:d8b4481a9d9118ab3db9d45d769e00ef91e91fa10a64f174648c16f5ac68946d

Observation faad95cb-ed2f-4b8d-8d8c-e712d49709e3 · outbound

This paper cites Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.373999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.373999Z digest=sha256:9e98995057841a420f0ceccd75860a34645354a8c4d586c3a4d3a0bce8de106b

Observation 0f8d7c8f-0779-4c12-b2e9-e5a4d6623cd4 · outbound

This paper cites VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:05:27.710092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:05:27.379194Z digest=sha256:8f142dc0afa855eaaae3e022e9244125937085480b2acd684e62f41bd11ad455

Observation 350d20ca-4483-415d-a886-5c9fe983c1ac · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:05:27.756352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T23:05:27.417239Z digest=sha256:c375884c80c084b2fce20d4e05399b9975a42754383962a41bfeb0aeab3d71fa

Observation a22d6344-4e16-45de-85ab-90bfdaa21930 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.403462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.403462Z digest=sha256:f2654ddc242438fb2a0f7e966370ed18c38c27257e35bc4fd4d7b71ed99e6e81

Observation 2defa89d-753a-45e7-a028-b0f7c1ae5177 · outbound

This paper cites Simple and Controllable Music Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Simple and Controllable Music Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.384167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.384167Z digest=sha256:6eba04d272d4385770d0393da575b10ca875361c10eabf3ca9e579c05a12032b

Observation 6cdbb547-8170-43c3-9e1e-9e21ee381a27 · outbound

This paper cites MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.388700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.388700Z digest=sha256:25047f179bfc11d85d3478ead6eb5c35536b0d8735c0b1f2510e61cfb297f72c

Pith citing papers

Observation dea83d22-1d9d-49da-ac40-e2eabc4cca0e · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.713763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.713763Z digest=sha256:00705d0d818ea4316aa3660228d9420bc19f8609a7d3b7a904549fe25cb0cfd5

Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.383496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.383496Z digest=sha256:497cbe38f28639b429beeddfea3079d118431f915f9f1dd48c7cda2ff8e6ca51

Observation d7a23a1c-972f-4725-8ad4-3706843219ac · inbound

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? cites this paper.

VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:48:34.544967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T21:46:43.353305Z digest=sha256:f50c4c09d42382d22c9ee8efb9631b1ba544319ad18bed1bfb2c063c99b28fc2

Observation cfb1d7bd-7733-41fe-aa49-4ab177adeeae · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.933593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:5de54dcfc8cf076aad43e6f536cf730e12198204be43cbe5623e35c1a8a37ea4

Observation 2d635f72-789b-474c-9e74-2db868550dc6 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.274194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:8344517bfada49a52b2150b204e424ba79c3e406c206208cc60f3bc4df154f16

Observation e0f11733-31b9-4529-adee-38832a8e15ee · inbound

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling cites this paper.

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:06:14.373792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T06:44:36.000353Z digest=sha256:c6d6d25dbc41db4894fcf26a2e6f4378dd9d58f734806b471a3c645f87ff80e3

Observation 1574171a-d19c-4fb8-a268-7320017c7ec7 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:45.368384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:de06ea606b4d5fafbe5525ba67c864868f22dae914c987ccb1ccd92701500dd4

Observation f4d24b73-9292-4f2d-87f7-8cf8abc57f52 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.852450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:b23b847b3034b232b7a0e11bf5b8bc82ae8a86c764926e904a4e47f86774c7c3

Observation 01c6ea4a-f099-458b-8219-85773421ecd2 · inbound

Inference-Time Scaling for Joint Audio-Video Generation cites this paper.

Inference-Time Scaling for Joint Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:40.932702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:49:20.194990Z digest=sha256:7f57fe026c3b3285c78ce7ab2c68c0cfbce87d3dfb4557ba6e457ab7f746908f

Observation 00a524b9-2b2f-4888-88d2-c28c0e276ec6 · inbound

AcoustiTrace: When Plausible Sound Violates Physics cites this paper.

AcoustiTrace: When Plausible Sound Violates Physics SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:07:18.202024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:07:18.202024Z digest=sha256:8048e96841267cbb4ee52bd4e354885db3e5f56d5e61fe618e89782dab789770