Pith. sign in

Paper Citation Record · LEDGER

Wan-S2V: Audio-Driven Cinematic Video Generation

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 32 inbound Pith citation observations for arXiv:2508.18621.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18621 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:24:59.444766Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:05:47.808077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.703864Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4968f6d7-9c97-4a4f-b775-bba1ef82af99 · outbound

This paper cites Qwen2.5-VL Technical Report.

Wan-S2V: Audio-Driven Cinematic Video Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.584884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.584884Z digest=sha256:dbcd7650fc2c032420696206caaf0bd9f6c0cfc96b08649aee8ce606b8674a31

Observation 7fb76f94-98a9-4e52-b57a-a60ac840fe62 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.824740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.824740Z digest=sha256:e6bce49aa3a6850aef6ecfa1f9bc12b9fd8c654c2d35a3ccb822735371819180

Observation 04bda69f-6eb3-4468-bb05-9d351ea780ed · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Wan-S2V: Audio-Driven Cinematic Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.874741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.874741Z digest=sha256:c36d600cceca31871e68c00d3aa1785378d6f6603d193921239ba3fb5028eeca

Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.025790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.025790Z digest=sha256:96d98db0b32d68ef60319930fc39284808e402aff558b623f9e50a05b9700bbb

Observation 3ae816ba-4882-42a5-9c5f-4b5e41715aa9 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Wan-S2V: Audio-Driven Cinematic Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.084742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.084742Z digest=sha256:19c81dffbc3aae5ad3c35811cef8ff5622a51834beb99222ede54ca909ac99b7

Observation 54d8d8a5-a960-4da1-b302-d210086694cd · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Wan-S2V: Audio-Driven Cinematic Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.254749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.254749Z digest=sha256:26332e4236c429cfd9cbb3f2d6bb74e1b38ab898d62573c87c1bb02ea4822e48

Observation 163e6c2f-8abb-466f-b17d-b21ee7be9c84 · outbound

This paper cites Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content.

Wan-S2V: Audio-Driven Cinematic Video Generation Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.301375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.301375Z digest=sha256:5dfebbcd3d3b419f573c9e5804b6b583777fd8d029c027a2f753bdde467a93d1

Observation 0ac0e266-00bd-4c76-9f87-ba7e83315f05 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

Wan-S2V: Audio-Driven Cinematic Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.444766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.444766Z digest=sha256:3dc7e8a14a86e1b723879c3b94954bc0a107d61d226dbb23ffd85dbf0f80acd3

Observation 13af52ce-5fa9-44ae-8acb-0f90b09b0c9c · outbound

This paper cites Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin.

Wan-S2V: Audio-Driven Cinematic Video Generation Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou Hou, Annan Wang, Wenxiu Sun Sun, Qiong Yan, and Weisi Lin

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.354746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.354746Z digest=sha256:c8464a6bd74958d787e253f8d309a5796635588568ac4835697ae4d6f2f62b51

Observation ca1bd54f-4053-41a3-b5f2-245a5bca4ded · outbound

This paper cites an unresolved cited work.

Wan-S2V: Audio-Driven Cinematic Video Generation Unresolved cited work

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.774990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.774990Z digest=sha256:a7d97c9c6d1cd88c9322b50d7ba3238a426733dea4df028070b7a84c1d247146

Observation b7516b7e-1033-429c-b33c-0060dd4bfa65 · outbound

This paper cites Christoph Schuhmann.

Wan-S2V: Audio-Driven Cinematic Video Generation Christoph Schuhmann

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.974740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.974740Z digest=sha256:0003848099ff3b97c42c57aaabe68f46ccf821c4ccd89f4c8ab3e517e17118bb

Observation b29b68cb-f0fc-44d7-9b16-1208154a8507 · outbound

This paper cites Image quality metrics: Psnr vs.

Wan-S2V: Audio-Driven Cinematic Video Generation Image quality metrics: Psnr vs

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:25:02.079683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:24:58.724742Z digest=sha256:25d55d5c012bb621d363b4c9fd139c9c2c0b4bb6566d4e7fed21b53e1130620b

Observation d77ea705-6962-4e0f-b2a1-7babb4f8852b · outbound

This paper cites ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation.

Wan-S2V: Audio-Driven Cinematic Video Generation ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:59.404743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:59.404743Z digest=sha256:8f1ee5ec97a9f73ad487c757d0cd96b5cfe96d6f067e61a28fb330109b7c5dc6

Observation 90c58207-0cbf-4150-bb93-e2f38d6b30c1 · outbound

This paper cites Flow Matching for Generative Modeling.

Wan-S2V: Audio-Driven Cinematic Video Generation Flow Matching for Generative Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.924743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.924743Z digest=sha256:1651e85eae7f6e05a01e6a41eb2ca5a04169357b0ec80c9dec1425666bffb9f9

Observation d377db72-a42b-49dc-a21b-bcd2603004dc · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

Wan-S2V: Audio-Driven Cinematic Video Generation USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.674741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.674741Z digest=sha256:71c5d1bdfa7403261ab3427f15918a2684ff13942e2ab8c10a48636e24ec6d5d

Observation 3e091873-51f3-4079-b666-ccbce3d342bf · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Wan-S2V: Audio-Driven Cinematic Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:58.625126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:24:58.625126Z digest=sha256:2c1be9e05fa3d77217958b2e5e6e43ee05adcf443229be887cf142727f6eb8b1

Pith citing papers

Observation 823551b6-d303-469f-a713-b7cff82979e0 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.808077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.808077Z digest=sha256:998b2ecd1212653fb11c9c8f5d248e0e872026aceebe6a01cad3528982b5cb36

Observation 31fba0ea-5d3d-4267-8785-c127026cd2ba · inbound

ASTRA: Let Arbitrary Subjects Transform in Video Editing cites this paper.

ASTRA: Let Arbitrary Subjects Transform in Video Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:16:14.277804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:13:10.141426Z digest=sha256:1f3a4945b44869f0892869ec7153032a282152e6229783b2252659e492ac9b01

Observation 6e6f1077-bff7-42c6-91f1-b22c9c3940db · inbound

Understanding, Accelerating, and Improving MeanFlow Training cites this paper.

Understanding, Accelerating, and Improving MeanFlow Training Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:39:16.078524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:39:16.078524Z digest=sha256:c05c05ba4813df48201737d8d141941c139e5d5240643a7e9d70276b04c3b672

Observation d3775d6b-28f1-4e2d-8dd4-07c78e054cc5 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.463818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:f83d79e1dd25697c668762e301f2fe95b4624c30b5fa312dd2056fe023f2220f

Observation 81c585a6-fcc6-41ec-8429-171803c4fc2b · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:41.295310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:41.295310Z digest=sha256:81d8244dd3464c50fc619191690a0369fa4a594791e3e551e8879561196eb2cf

Observation f1e3d810-2745-4a83-8ad8-805a95a82067 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.546710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.546710Z digest=sha256:915b37fbe5b519ca64ccbe3f5fcbd51443eaf18def0a655929cef7a84d3fb4b7

Observation c39098de-37ea-4f0a-9dee-cc030938ab7f · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:06:20.610336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:a362d450b7c67759e67cc9d478ee586e9561a2d497298106f4622a336ed395ab

Observation aeee2f62-f7d9-4b5a-93e4-43ba92da5b75 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.238327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:c9bef44f033f2bc0750fb9bc29d1f97c2ff81138f444d8601f1cc97607671849

Observation b2b5bc18-be68-4dac-98bf-d3478a9c8bab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.182089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:fb69f9c26757bf4e1e5c363354d27243d6041c8bfcf2f2bdc6e425be9fefffda

Observation ed9b6ca1-feca-4e93-8d37-466dc6ccce21 · inbound

LPM 1.0: Video-based Character Performance Model cites this paper.

LPM 1.0: Video-based Character Performance Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:30:59.901752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:09:17.360697Z digest=sha256:6dddfb24e68c9b99d3764fd49ab0e789568abb6f58b467b2c46ac9dd0577f10c

Observation 53a221fe-6b10-4d3c-91ec-02fc5e31bc78 · inbound

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence cites this paper.

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.586177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:13:06.005185Z digest=sha256:f9b44577af980972619b687964935528fe545f50a6822e0d7e7ff8ffec541bb6

Observation 915f1fcc-8e73-4951-b8ff-7eabc64a9178 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:04.221942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:abed929e99c268fcb4d2ccef3b9303be10ca883f6ff3a6fe38088c528027d375

Observation 3ea3c7b8-d52c-47e6-9c48-0973c96485da · inbound

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation cites this paper.

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:14.635424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:54:23.108142Z digest=sha256:08eadf36ff53ac5fcb6ea103414fc24906e61adf53e10438b65f511aafc7d340

Observation 0615e64f-6599-41bd-83d8-5bcb33534e7c · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.776769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:dd1d996b3c27fea55421bfdb4122dcd75f35ae58549832dac44229c232e6415a

Observation a3dc51c6-bb87-40c8-b871-35b2a91dfdc1 · inbound

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields cites this paper.

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:06.456958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:50:37.638842Z digest=sha256:968804162e85739fda50cf0ed4f2b197cb94ff41486428f10a8a91f87e66b122

Observation 37293b50-a45c-4a27-a122-eaa91e372b17 · inbound

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors cites this paper.

VidSplat: Gaussian Splatting Reconstruction with Geometry-Guided Video Diffusion Priors Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:06.244777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:00:29.091220Z digest=sha256:df589e2361f0cb63de49a7d608df95150f91308efd47de426d4aa3ec53eb2f73

Observation 85ee8bdd-7e3f-4f54-b055-863779596bf1 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.362381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:0e5998c043b0b76e8fd89c7739813b522789489890fed0fbf56630ba4a7e7708

Observation fe6b0a7c-2970-472b-a76c-05962dcf8ef6 · inbound

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration cites this paper.

StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:24:00.909558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:15:17.804936Z digest=sha256:f469e62245230f679e85068fc4e959b49477a289c69e866c9a43bd8e43516be9

Observation 39e56e96-9933-4318-b211-5bfd45fd9511 · inbound

LongCat-Video-Avatar 1.5 Technical Report cites this paper.

LongCat-Video-Avatar 1.5 Technical Report Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:50.607869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:28:16.968339Z digest=sha256:9079054663438468fe1ec9b5b2fdba22879b323f4a7bb34758db4aba139c0a0f

Observation 802dd96e-2a18-45e5-94a0-861643019afd · inbound

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance cites this paper.

ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech Guidance Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.705373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:56:35.931176Z digest=sha256:41f6595d5d421da656ad6d85c426703b655f0c94a6eed239b2812ea802bd680d

Observation 9adfd245-c7ea-4461-bb45-c772ba35572c · inbound

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data cites this paper.

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:44:19.552832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:35:58.943585Z digest=sha256:8c1d8774f004fbf8b56b3c777721bec4511318e5cb655ad15616e27c59fe2891

Observation abda8adc-4997-457f-95ab-4eb69b4bc66c · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.338694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:db1294d6cf7b40d4ad088bc07d07554ab1052a069434db644bdb782cd5aa2270

Observation e55d4727-a3f8-4397-a63d-f09874dc90a8 · inbound

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models cites this paper.

Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:16:58.093129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T13:16:16.676244Z digest=sha256:2fac481a174447f90b2a23af2ed76df01df6069549ec8e57787dbfd6eaa635d8

Observation 6500ddfb-d75d-4892-a023-fe235c500c26 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T04:46:57.581402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:46:57.581402Z digest=sha256:c417630e9274797a30fe9dbe7fb29deaf5b18a863717a344952680c4d598a1dd

Observation f9925a27-de13-409c-9838-6f113ba0d756 · inbound

Vidu S1: A Real-Time Interactive Video Generation Model cites this paper.

Vidu S1: A Real-Time Interactive Video Generation Model Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:31.525012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:31.525012Z digest=sha256:e435f77321a9c1232db2fec26cb444c05c33cbf1795fd85e2ad70a4ed32ac1f9

Observation 8d80c2af-be36-458f-a467-b2f597e4d5ab · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T01:59:43.167178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:59:43.167178Z digest=sha256:90e09d0d901ab11828210d8e0eac6b113776eb16ae8331b2de2b953e6149eedb

Observation 96151756-fd25-4fe9-a567-e7ac1b8a8321 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:10:13.757131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:10:13.757131Z digest=sha256:7a0dda8f869fd99203269472977890a63fd2aacc1a6efbe30a98223a352cae7c

Observation 0db1233f-e7bb-4a10-bb85-3d0ff1a05a65 · inbound

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation cites this paper.

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T07:35:22.880155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:35:22.880155Z digest=sha256:5360e63d16ea9054ddbe61970804464444c11bcc4e6e101b58b50af91c12521f

Observation 51f7c75e-916d-4188-b3d9-db5c41a5be5a · inbound

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars cites this paper.

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:59.468371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:18:59.468371Z digest=sha256:57fa796575972a2fc84182b90d20e53bc67dffb0d414fbe174c4cc58621e8085

Observation 32ba9d0c-e283-4059-9009-9515cc4f2452 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:11.944416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:11.944416Z digest=sha256:3460daa611d2bba1dc205f1358ec9576b58f8c319623d23ce7ca6189e0aaf311

Observation cbc871b1-6b14-47db-b6b1-2c0f3c82aaef · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:13.104187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:13.104187Z digest=sha256:9095a374cc66617e9842d57f6a096d80d4004d21b8875b05c90165d3fea607e9

Observation afd2025c-a931-4640-a65e-51dd3c8b3371 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation Wan-S2V: Audio-Driven Cinematic Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:46.515828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:46.515828Z digest=sha256:88b1ae8055e9e6b2cf0fa698ea048480ce05c26cb22af19d9a1b31f711edf9eb