Pith. sign in

Paper Citation Record · LEDGER

Audio-Sync Video Generation with Multi-Stream Temporal Control

As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2506.08003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08003 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:24:30.637501Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T11:34:32.558440Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:38:14.727310Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59ac458d-94e1-44f3-8a0d-867e9db24c5e · outbound

This paper cites Diverse and aligned audio-to- video generation via text-to-video model adaptation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Diverse and aligned audio-to- video generation via text-to-video model adaptation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.824720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.255357Z digest=sha256:bdee2af6e99efbe4a05ebc6422a16c2df15b05d3700a3163f31ecf9bd84361ab

Observation 6e1fd37e-ec24-4972-92c5-1cf29c734613 · outbound

This paper cites Long video generation with time-agnostic vqgan and time-sensitive transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Long video generation with time-agnostic vqgan and time-sensitive transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.809045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.260509Z digest=sha256:9e20ec5845b3d8e30c03e8fbbe58f3746ed5428b695a17bb0e0b8265d3e831c1

Observation 6aa8241b-dfda-4759-b83b-3a53ba12a598 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.793603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.265481Z digest=sha256:20684497ecb8dc5a066947c4a975a3294225591dd1471e45a5c48ed85f034251

Observation 173c70ac-33a1-4244-a17b-a8251bd92aaa · outbound

This paper cites Sound-guided semantic video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sound-guided semantic video generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.778883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.270005Z digest=sha256:41fabb79a87d84c3c91cc5908a0ea7c289c6a115c1fb924dcbca3092f4f1a748

Observation 104fdeee-48a2-4e1b-8463-d7fcf09cdd8f · outbound

This paper cites MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.764037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.274528Z digest=sha256:56fd516d3ecb849291a1155584a172e48f76f439507fddcf08a9cf7e4dce8e8a

Observation eb9bff16-e028-4a01-bc6a-7395298b8204 · outbound

This paper cites TA2V: Text-audio guided video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control TA2V: Text-audio guided video generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.749032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.279142Z digest=sha256:b06bb7a12f67b9282f8fd79cac4d2de399b57f716f8008fa0de2111d174c83d3

Observation 02f5e58b-3e74-41af-9b4c-8f011ee5de61 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.732415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.284477Z digest=sha256:f89d56188dc310ab1195da1a01c2ac442f505a6e56b4659b632e449e538ec705

Observation 5b0f35db-ba58-4d9d-ba5f-624f0791daca · outbound

This paper cites Speech drives templates: Co-speech gesture synthesis with learned templates,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Speech drives templates: Co-speech gesture synthesis with learned templates,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.716351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.289145Z digest=sha256:dd7b7761cdc469d9632a10fea3f046932638a80c1a09a7288ffe5698a5088884

Observation 9db135e9-97a3-4b6e-9623-a31610f879cf · outbound

This paper cites Visualize music using generative arts,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Visualize music using generative arts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.700762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.293448Z digest=sha256:687c6901656562af9ff41a1378abc41b676930bd440054c2c92b0fcadc51c7fe

Observation b6fa7cb1-0213-4c01-84e5-2fb459e1f6a5 · outbound

This paper cites CogVideox: Text-to-video diffusion models with an expert transformer,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CogVideox: Text-to-video diffusion models with an expert transformer,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.684858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.298059Z digest=sha256:fc550491825ee74278eafbe9e6b37851261bd531d762bd32766254c745e0ac0a

Observation 53b57413-2458-448f-9c31-6ddfcba84fd7 · outbound

This paper cites Structure and content- guided video synthesis with diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Structure and content- guided video synthesis with diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.669493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.302949Z digest=sha256:218712c35f16017c267ff634a493c9474f4454b542e9cc256d6335271c96e364

Observation 44cb9eee-3cfe-4393-bca8-532eafaafac4 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Audio-Sync Video Generation with Multi-Stream Temporal Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.307892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.307892Z digest=sha256:713d656257426ee82e8698dfc83590e82f9d6f50bfa4b53b3b8819b01ce1e0ea

Observation 90957916-2290-4432-99d0-5096cb5e939f · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.313733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.313733Z digest=sha256:8d1182a997f1eb3ac0dcf5db6322a473286c035588948c4886763bd93c8daf12

Observation e1db0a67-136b-46e0-98a4-1d8f3ddb1577 · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control High-resolution image synthesis with latent diffusion models,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.318712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.318712Z digest=sha256:ae250e1d3abb23a60d242bc68e0558be98982da21654d7027bd4b009fa244196

Observation 130049ee-dfac-44ae-9cf8-72b93248efbc · outbound

This paper cites Interpretable 3D human action analysis with temporal convolutional networks,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Interpretable 3D human action analysis with temporal convolutional networks,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.641049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.324986Z digest=sha256:2a2a50ff9ec635d6e744651c32d93c168fb5a97f56b14fedfe1864175ff4c33b

Observation 9fcda4e2-2151-4ea0-bb10-0a8fbb1ac67c · outbound

This paper cites Is space-time attention all you need for video understanding?,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Is space-time attention all you need for video understanding?,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.625063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.330694Z digest=sha256:9ba854b53b7b7caf2e14d1f1aab62d4b30531dc925bfac1b7c816dcae48eead8

Observation 5086d922-64e6-440e-be42-6440bf04d4c8 · outbound

This paper cites U-Net: Convolutional networks for biomedical image segmentation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control U-Net: Convolutional networks for biomedical image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.606277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.335384Z digest=sha256:182adc4ff772fe51054a0f7c9aaf52a07fcca98cb9333c400703c7349e187752

Observation af590379-9528-40cd-a146-62c41fa36b46 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.341073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.341073Z digest=sha256:e0779ee00b1378e4a360026809f06e90544d0d57c8e538fc5f3b630d0468ac8f

Observation 4cb69949-eeaa-4b83-976e-1315e82e3419 · outbound

This paper cites Scalable diffusion models with transformers,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Scalable diffusion models with transformers,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.349078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.349078Z digest=sha256:6a05be407f4cc43bec552fea68fde6a6c5c740f284a256751e4fc0f154ad0a01

Observation f185d2ab-9d88-4072-bcac-028ccd651926 · outbound

This paper cites Language model beats diffusion – tokenizer is key to visual generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Language model beats diffusion – tokenizer is key to visual generation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.579163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.354412Z digest=sha256:bbef9c3365ee53fd537b5e06941d01b59c41ac39a57aff63fd68c7d6b5584321

Observation 74cacca0-99d8-497b-a584-3fae50b1c88e · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control Wan: Open and Advanced Large-Scale Video Generative Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.363858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.363858Z digest=sha256:c19f7a7ce946ad4150f3fea63f8a298d792d463a953038129a38b2947bc2721f

Observation c90e47a0-d7d3-4a13-b202-f45f984da477 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.370136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.370136Z digest=sha256:940c1aa518d8609d6bb2280c9fa1cc0dbb9bf25878f5a3a1e09180ccc2492888

Observation 1c9c5c64-36ce-4a6d-9ade-88715a693c0a · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

Audio-Sync Video Generation with Multi-Stream Temporal Control Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.377023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.377023Z digest=sha256:d6e828a2fb3e425bff4704176010f9af944e1b4f44f7a5c810a2cf4135cf3787

Observation 9386c427-f05d-4b6c-9596-cf0abfce3c4f · outbound

This paper cites Sound2Sight: Generating visual dynamics from sound and context,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Sound2Sight: Generating visual dynamics from sound and context,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.564156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.382367Z digest=sha256:936a0216c19cf7e783cf7fd03bdb48c7557fb7b7e703661989139d0524463046

Observation 482fc39c-24fe-4b38-896d-706d3e01cbd0 · outbound

This paper cites CCVS: Context-aware controllable video synthesis,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CCVS: Context-aware controllable video synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.549939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.387065Z digest=sha256:d19c9d577d4e82c22625d1f99dd8853889ba2b45d235599bbe6e95ebf1a9f5e9

Observation 31c2da8d-43a4-4bf0-a56a-2187c03388e3 · outbound

This paper cites The power of sound (TPoS): Audio reactive video generation with stable diffusion,.

Audio-Sync Video Generation with Multi-Stream Temporal Control The power of sound (TPoS): Audio reactive video generation with stable diffusion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.535252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.392168Z digest=sha256:1822397897113b3745849326bb98f3ef46a61b0ec9d77c3652cd7ca502a6a295

Observation 0da452e7-443d-444a-b16d-2020746276fd · outbound

This paper cites Audio-synchronized visual animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Audio-synchronized visual animation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.520802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.397425Z digest=sha256:e36d16191e984633d587a8ce1a28e64f93f3177a70a169080783c06d9e58c70b

Observation d141dff9-2d35-4719-a13c-b94753f9f633 · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Loopy: Taming audio-driven portrait avatar with long-term motion dependency,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.505797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.402130Z digest=sha256:42a51dc22b442dcf6c9e928d4e0b99c55ac27e5c3d38ee43d9e78a223de7895e

Observation c3a1929b-d64c-4a7e-b2a7-33342cbf61a2 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Audio-Sync Video Generation with Multi-Stream Temporal Control AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.408110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.408110Z digest=sha256:ae5633ea4cd2c02538fb3e68ef9afe53b11d4695083cc2ee92a5f30874dc8155

Observation 7fe47a68-6801-4dd5-bcc5-7f8468080216 · outbound

This paper cites CyberHost: A one-stage diffusion framework for audio-driven talking body generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CyberHost: A one-stage diffusion framework for audio-driven talking body generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.476293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.419495Z digest=sha256:8cae50218e44d3af68c5d842671370b7f658a174a0a52e02e0d9a23ddee9f188

Observation 03919a94-3a27-4efb-87cc-cd6bf72d9b33 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.424440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.424440Z digest=sha256:021d632a08a2954af103602965ed85732c7bec08b93d282cd7c24e0fab3ac1d1

Observation 97c70973-2b49-4825-8c6e-fa7e72169ac0 · outbound

This paper cites Dance any beat: Blending beats with visuals in dance video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Dance any beat: Blending beats with visuals in dance video generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.461031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.429761Z digest=sha256:c81fbc8738bb87f97aa2baf2717a42e8acfd65fb3d6c9f12d1dc7173740cc929

Observation c4d9c21c-b91b-4105-94ef-c6ca6d463b0f · outbound

This paper cites X-Dancer: Expressive Music to Human Dance Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control X-Dancer: Expressive Music to Human Dance Video Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.435331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.435331Z digest=sha256:40a92e03e4afc48aaba889be3c8ee46091716d954fc932e046485dfdd44f3e06

Observation 7a10a7d8-b6b4-48a6-9fbc-5f009fdd31da · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Taming transformers for high-resolution image synthesis,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.444999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.442078Z digest=sha256:7eca6fe6038d87604faf5fd40063f36491d7908b7a872c32f022313f4a679e54

Observation dd210d7b-cf20-478e-bc0d-459bbdea9649 · outbound

This paper cites A style-based generator architecture for generative adversarial networks,.

Audio-Sync Video Generation with Multi-Stream Temporal Control A style-based generator architecture for generative adversarial networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.428879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.447252Z digest=sha256:58aea6e8daadc836afe2dab6f2b2080dd94ac478693163cfbed695292aea6f4c

Observation 0292bced-28bf-4495-94c8-8a5709d65e6f · outbound

This paper cites Tr\"aumerAI: Dreaming Music with StyleGAN.

Audio-Sync Video Generation with Multi-Stream Temporal Control Tr\"aumerAI: Dreaming Music with StyleGAN

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.452897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.452897Z digest=sha256:c6d3954127ec88a92596b495ff4f57f8a47ab96888091069245c2b118dea48af

Observation 0a69a0f0-b5f1-4d9e-8971-7ff74ac49f2d · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Audio-Sync Video Generation with Multi-Stream Temporal Control UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.458514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.458514Z digest=sha256:e0c48518829b8ebee46c7bc0e1017383601c07ce2f9d2f4831ab172c1b51ba2c

Observation c80b6ac1-61f6-41c0-b13c-5c9b8137ce49 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Audio-Sync Video Generation with Multi-Stream Temporal Control Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.463937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.463937Z digest=sha256:1424e140dce747159db40286b7e1d3e6af081846967814423519fabd9744d910

Observation a8e9d9d0-2234-436d-b21d-f7333d88753a · outbound

This paper cites AudioSet: An ontology and human-labeled dataset for audio events,.

Audio-Sync Video Generation with Multi-Stream Temporal Control AudioSet: An ontology and human-labeled dataset for audio events,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.413510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.469091Z digest=sha256:7111bd0f1e3787175c7123f4973b0e5983a3bc3108a25c1cf5b9d0f22c684d94

Observation d1a3ecbc-0034-46f2-bdcb-4a9efbaef075 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Audio-Sync Video Generation with Multi-Stream Temporal Control VoxCeleb2: Deep Speaker Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.474142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.474142Z digest=sha256:7d2c4333874cc6a77c7bc03cfb923f67f67dc5c3dde6714e3ba32a127bb5bfb4

Observation 7b68e6ec-ae47-4859-815e-e0e7d579afca · outbound

This paper cites VggSound: A large-scale audio-visual dataset,.

Audio-Sync Video Generation with Multi-Stream Temporal Control VggSound: A large-scale audio-visual dataset,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.394335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.479131Z digest=sha256:ef43d38df62948063da04b8d64e980e0157fc78fc6a27e0819deed9386730dfa

Observation 6bfa8e3d-ce63-4a43-983c-962631ef68ff · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.378605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.484063Z digest=sha256:6e8453dcdd755e3cdefeb0bc5176b5e0e23637ef539069b5e20dd78044114270

Observation 32fdc992-0e22-406b-ab5b-a58d40724d58 · outbound

This paper cites InternVid: A large-scale video-text dataset for multimodal understanding and generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control InternVid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.363011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.488419Z digest=sha256:b5181dc47e2e8d6ce10058ac18344e709cdc06be8c3f3d02612172beef278481

Observation a7f64bdf-cd4f-43d4-848f-6e33548caf24 · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset,.

Audio-Sync Video Generation with Multi-Stream Temporal Control CelebV-HQ: A large-scale video facial attributes dataset,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.492770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.492770Z digest=sha256:147971bf2429d3dbc461c299284b15c89f431b882a44f24a4dc5b0c31663a943

Observation 02f33fb4-2f3b-4c8c-8b03-06a0ea14503c · outbound

This paper cites MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.497996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.497996Z digest=sha256:cc18a74f4d37c91df87f2cb95088c3ee0683343515c5b41749946b713d223540

Observation 87cd5216-68e7-4058-b722-039abb8eaf52 · outbound

This paper cites Condensed movies: Story based retrieval with contextual embeddings,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Condensed movies: Story based retrieval with contextual embeddings,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.336609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.503180Z digest=sha256:f9aa029adea7e268cebee68365b4e6b076083cdbedc912964d3a50148495128e

Observation 8e73c803-64ed-40bb-a691-b58d78a20ef3 · outbound

This paper cites Long Story Short: Story-level Video Understanding from 20K Short Films.

Audio-Sync Video Generation with Multi-Stream Temporal Control Long Story Short: Story-level Video Understanding from 20K Short Films

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.507687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.507687Z digest=sha256:2e75b1b53de51adef43ecc8aa0f4d439a96338a222ecd7854fea3c4a90a3230c

Observation dcc8e9f4-e5a2-46f9-8abb-7876849b57ea · outbound

This paper cites VideoCrafter2: Over- coming data limitations for high-quality video diffusion models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCrafter2: Over- coming data limitations for high-quality video diffusion models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.319998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.512852Z digest=sha256:34a79b9cc64b489bd2a65c02d32082613b604a6b964bdfe4d99e8bcca6bed7a1

Observation d62e1bd4-120d-4c6d-9a71-f055e92dacbb · outbound

This paper cites Video cut detection and analysis tool.

Audio-Sync Video Generation with Multi-Stream Temporal Control Video cut detection and analysis tool

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.304988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.517286Z digest=sha256:f7e0be8f7a60e4194db44c1ab6ef28bc0c2081919fd048277c26c382f8641125

Observation b8fdad6b-7583-4436-a9d0-30ad09f57a16 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

Audio-Sync Video Generation with Multi-Stream Temporal Control Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.522443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.522443Z digest=sha256:0ff1e28183024d1f16d3b9f3020c747f03639f384a353669eaaeda7eaa2bdadf

Observation 5a25c88f-a587-46d3-8cd0-b7f0467bacc4 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Audio-Sync Video Generation with Multi-Stream Temporal Control LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.527178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.527178Z digest=sha256:3169504028a04e20c58a2e11a8294f1c6ad1d5e81b1435cc18b425f64dcc3bc5

Observation 2fea16d8-7016-45c3-b221-8c19d508a5a2 · outbound

This paper cites Cinematic sound demixing.

Audio-Sync Video Generation with Multi-Stream Temporal Control Cinematic sound demixing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.289568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.532116Z digest=sha256:8ef71707f291388d1749b9967d4f50a4da026ba317cd7d562d21baf1b67b1a21

Observation d017bcb6-ff74-45fb-8a8a-c1960fa909e5 · outbound

This paper cites Spleeter: a fast and efficient music source separation tool with pre-trained models,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Spleeter: a fast and efficient music source separation tool with pre-trained models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.273725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.536937Z digest=sha256:6cf95412a3f14c70fcd9a28862e4f290e87b5a48a03a9e2a7043965b69d1a1e0

Observation 11c5ce5e-efa2-41ea-8d26-e62ee7767c38 · outbound

This paper cites Ultralytics YOLO.

Audio-Sync Video Generation with Multi-Stream Temporal Control Ultralytics YOLO

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.257250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.541280Z digest=sha256:2ce5978a864da41ba3d8f0b21f7acf0e108e1f1f3e6e5397c8bf29394015ba21

Observation fe301372-756c-4a2f-905c-8cc24c4c07ce · outbound

This paper cites Meet scribe.

Audio-Sync Video Generation with Multi-Stream Temporal Control Meet scribe

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.242082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.546019Z digest=sha256:af437a4e3dc2e03937eeba6dc0639b85e39dbdb2f8cd1efb43e4338d4d16cb12

Observation c24d23f4-aca3-4d3f-a2ef-5d519b718316 · outbound

This paper cites TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction.

Audio-Sync Video Generation with Multi-Stream Temporal Control TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.551203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.551203Z digest=sha256:17337d03088f48eb34f09381d6089d04813a46bc0285128e63acfe7cdafaa462

Observation 0505b90b-5a1f-4c55-9af1-34f06ae77fe3 · outbound

This paper cites wav2vec 2.0: a framework for self-supervised learning of speech representations,.

Audio-Sync Video Generation with Multi-Stream Temporal Control wav2vec 2.0: a framework for self-supervised learning of speech representations,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.225434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.556119Z digest=sha256:7c46c0cda8c0a1592ccb4b3b6fcc9c004f8e795f09cf4033289442b3a84b147c

Observation de1ebc8c-a82f-4075-9e3f-21ffa472914a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Audio-Sync Video Generation with Multi-Stream Temporal Control Adam: A Method for Stochastic Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.561313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.561313Z digest=sha256:fecce7fe4c97cd0f25d0f60fd049d862713a794c43bb4326b5a5f17181c8d453

Observation 7d5800e6-a8a9-4933-8e39-85d99e5040c8 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Audio-Sync Video Generation with Multi-Stream Temporal Control Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.566199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.566199Z digest=sha256:ee67805559cef7272d29b8a36d38810ca39d8c593232600f54fde7d2e543b09a

Observation 861757ed-d16a-4659-8865-6d48a251e24f · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Learning transferable visual models from natural language supervi- sion,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.207263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.571380Z digest=sha256:ea1a4e4db5e0656206b57a81d845bdc04beaf9ba20730ac867096463606b1108

Observation ccb49950-47e0-4796-a667-e079589500d6 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.577220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.577220Z digest=sha256:8e378de1ac39921073251f2ade11b200a42fabda0f424aa91f248ffe1044fb3d

Observation fbd91222-ad01-4fd6-9b32-3c82239582fe · outbound

This paper cites ImageBind: One embedding space to bind them all,.

Audio-Sync Video Generation with Multi-Stream Temporal Control ImageBind: One embedding space to bind them all,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.190770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.582083Z digest=sha256:002a838e58dfe64cd7d5b02c2cd5a03a316fc3959ff8a8ec2126486de1d0feb4

Observation 646ecfe8-29ef-4aba-af15-a3ad89219233 · outbound

This paper cites Out of time: automated lip sync in the wild,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Out of time: automated lip sync in the wild,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.175016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.587157Z digest=sha256:da07aec758f33af5b0fffd71a57ba3456d2fdf93bffca5b33c6f76d227bcab14

Observation b9b7134a-a243-48a1-95e5-66c7ac1cf334 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.598574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.598574Z digest=sha256:62e7ec88c0ca083a2e91349ed34e9390b065bffe011f75e32bfbd541e025847d

Observation 2f578fe6-ae65-46dc-8359-c8063755c4de · outbound

This paper cites TA VGBench: Benchmarking text to audible-video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control TA VGBench: Benchmarking text to audible-video generation,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.157131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.604212Z digest=sha256:236bf43ca6305c4ed13d7c927c7e4048f3ac7ce590f9cdd5bdd3baf9fc010cae

Observation 1e644474-6768-4be7-b191-5aef89235a09 · outbound

This paper cites MMDisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control MMDisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.137563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.610563Z digest=sha256:0be075067513701dcb74a5080311f341952b7e147465d6924cce524496bb42f9

Observation d3aa9b78-8bee-4c6f-a667-ef4b871127dc · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

Audio-Sync Video Generation with Multi-Stream Temporal Control AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.615859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.615859Z digest=sha256:597b596a39c0f420c13ade98de3f4f42cbf09d31f1eeecfe0210ca17000d5de6

Observation c59b8c40-6eb5-4122-870f-29532280bded · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait image animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo2: Long-duration and high-resolution audio-driven portrait image animation,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.621496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.621496Z digest=sha256:00f461d44db61233c88b2560c6b632a0f9bb80c33feff61b6904f1b8befd00ac

Observation 21c9b3e6-2ec4-4247-8fa4-8220abff7b6f · outbound

This paper cites SadTalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,.

Audio-Sync Video Generation with Multi-Stream Temporal Control SadTalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.490873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.626811Z digest=sha256:7c0262a154c58f7e9ad7aa1fb5e23979b378a9d75d2d5655950602e4bac2c4dd

Observation e04b7e44-6576-4e27-a303-72d88c1d9e66 · outbound

This paper cites Maximum filter vibrato suppression for onset detection,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Maximum filter vibrato suppression for onset detection,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.108216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.632022Z digest=sha256:904029b290f01ffa88a9594387c8d1d2d7c7d34ca6e465dfb77263c3b2c40246

Observation 533130be-c7b5-4938-a734-f57842d3dd34 · outbound

This paper cites Determining optical flow,.

Audio-Sync Video Generation with Multi-Stream Temporal Control Determining optical flow,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:24:31.087789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:24:30.637501Z digest=sha256:75f6620032b852b89801dcd184e6a6bd98d518db8cf8f066ae8e32b64737abba

Pith citing papers

Observation dea592b8-9aa9-4846-b81b-926220787e57 · inbound

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing cites this paper.

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing Audio-Sync Video Generation with Multi-Stream Temporal Control

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:38:14.728832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T11:34:32.558440Z digest=sha256:8dd004dce74fdd8596d3470d1b87bcf18d17a6431026c71594ae6c12e6f9813e