Pith. sign in

Paper Citation Record · LEDGER

SkyReels-A2: Compose Anything in Video Diffusion Transformers

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2504.02436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.02436 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:59:29.507056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.697793Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 39046ad6-2539-40e1-b5a4-c817f2aee04a · inbound

SkyReels-V2: Infinite-length Film Generative Model cites this paper.

SkyReels-V2: Infinite-length Film Generative Model SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:23:04.222693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:23:04.022599Z digest=sha256:39a3e8825f9d4561eed4bbb8d79adcf40f720102391b89c1ad358036d32ab0f4

Observation 87b78b0f-fefc-47bf-a394-e5f5bcd6680b · inbound

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation cites this paper.

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:59:29.507056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:59:29.507056Z digest=sha256:34b57f2733ec9047953884324a7eb0ef119a56a4a62a53d3748cb29f7ace2956

Observation ab65c0ea-f8a6-4c34-b126-694ca2bb456c · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:05.579680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:05.579680Z digest=sha256:c44f7303c6f7217a20441b321a18fc916247f6537164fc9a13ad0e323c792665

Observation d13d751c-5cc6-4351-86f9-1e9a92164b10 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.911278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.911278Z digest=sha256:7971fda39e8d26224385a0c82d2f5f34c2dd645ef809ee6584f5680e18c729b5

Observation 3578d72a-dae4-4e0f-9541-0faced2c218e · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:08.358935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:08.358935Z digest=sha256:e7ceb9f0ba37e88fe9b0ce331862a5bbdca3dfe3d12551df4a03fe1d4242a5a0

Observation d6acd82e-d179-47d7-9f00-b4b7cc9edfba · inbound

CustomX: Unified Character, Action, and Scene Customization in Video World Models cites this paper.

CustomX: Unified Character, Action, and Scene Customization in Video World Models SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:49.621536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:49.621536Z digest=sha256:fe60cbb33f9b80614619816b3ea6fc5e2d174c2931392eb552234baf1e496632

Observation 1b6d3fdb-d546-415e-8772-7bc04b4a3446 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:11.237810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:11.237810Z digest=sha256:009024e9e4792783b4ad977f4717b09cd9840045ca91cc8d8e390c32365d5479

Observation b1f8a611-a649-4508-8766-71186b1dff45 · inbound

RefAlign: Representation Alignment for Reference-to-Video Generation cites this paper.

RefAlign: Representation Alignment for Reference-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T18:01:19.034578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:01:19.034578Z digest=sha256:1894cfac5d55ae972a4c671386f4ebc8948a034245acf3f0912a0f854cf5fc86

Observation 2fc26fd8-3a6e-41f9-848c-bfa68a93da58 · inbound

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation cites this paper.

Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T12:23:05.876881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:23:05.876881Z digest=sha256:3dc3fdb2698bed78046db32b1484a78bb204fa9a557bda5b6f4758a94b5d755b

Observation 72e0e95b-1a83-4e3c-9ff6-e486b320c67c · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.736553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:befd8a9f53f4f6a3e9765d55513d0668dd41030cdd1ff337f42080b1bb5fcf6b

Observation 6d430c7f-fd44-475e-80bd-e20e4a92d472 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.589537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7834e13ac2a8cec263b8dc06e3ec75b95a33a2832d25ffbb3bc9b638d2408f61

Observation abfaf6de-042a-4f2f-a06f-c4a18735ba4f · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:27.920254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:3ae803e83758b4785913d78e19490f80fabd98ae2c347a7860a7636054a3fb37

Observation f809b158-7281-40e1-a246-02951b9a356a · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.822714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:9fc27bd58381f2e456ddff4f4dff947bc62a5bc418d5ec0af762e85fa9a0426c

Observation c9566795-c710-47f4-a259-cb5f6d37f777 · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.513027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:59f83edb2bf86f199160a517caefa4379ca7784ca872614b99b65c835edd6acd

Observation 214c9d33-2862-44ea-a967-9c7f20332a3b · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.606468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:3b8d3f95d6fb65ce39fbc7d4998dc6353ee181a9debc2b1d48ae455f7f975c0f

Observation 9aa3b2cc-8c15-4372-9f55-8689061fc327 · inbound

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation cites this paper.

FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:09.728834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:40:00.224358Z digest=sha256:e706b987b51766c23d13bd67a2322f691a1afd12eed5b60d206657248aa39f57

Observation 34ff1e3e-f549-4184-b7b4-3404c3260b75 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.556258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T20:11:33.277349Z digest=sha256:7b2b31c8c3d94e0996172549e9b240f05d665cda87379a4e004c8b5ac70115f7

Observation 8923338d-5b4a-4ab9-a032-b86c1e9aa642 · inbound

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization cites this paper.

FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:25:00.575779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T19:23:07.044659Z digest=sha256:185e5a49b6ed8c2ff35b9a0a792bf86cf0a51d6e0c2b0e9dcf57b00c1063c855

Observation 2d1a34fd-fe81-4098-a5e2-3b1286bbf6a7 · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.376116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:0ac53f0571658d94283fc98f3d4d3ccf0d7df3921960617cea66ade837b4b266

Observation 49ff3ff5-4d70-4335-814f-50e9c371f5c2 · inbound

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation cites this paper.

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:31:09.909379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:30:18.961069Z digest=sha256:9332b19898b39798ffe11692d0018aff154ef8c77ad925dc4dade33513db65e6

Observation 641621f0-0387-4d94-b33e-9082869dd837 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:41:10.408596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:8924321e5373e66f2d473667bd32b931802550622b7019af92534df3be0e953e

Observation 4243e34e-c835-4773-91d9-83118c6f6ee1 · inbound

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation cites this paper.

Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.490012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:25:22.778550Z digest=sha256:15b008d76e1e1a26edddc24d4caacaa7011605e96fb0d7d041b404fc96e7a967

Observation 5283a361-62e2-45a8-b4eb-f8d9ed9e2e13 · inbound

Streaming Video Generation with Streaming Force Control cites this paper.

Streaming Video Generation with Streaming Force Control SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.380960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:14:32.465663Z digest=sha256:3720a5cb8c3d2193f60f3f38791360f33d0df336ff8e56a173ef179f360793b1

Observation 0afa77da-e7c3-4a5a-a63a-19e530758b4b · inbound

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation cites this paper.

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.951440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:46:56.871174Z digest=sha256:5e1ab810a8978e99cabb686d1476a17d6d082108743f67a236569371fa330f38

Observation 929fbd2a-cce8-4503-a2d1-03045acf5f06 · inbound

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation cites this paper.

CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.689770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:51:31.625416Z digest=sha256:244c5b4b9f1fa0102171faff238b0fbdd3400f8e8ed31b01a4952329b8446ea1

Observation 08d03e3d-dd47-4ca5-88ce-e75f00e6a344 · inbound

DramaDirector: Geometry-Guided Short Drama Generation cites this paper.

DramaDirector: Geometry-Guided Short Drama Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:59:55.923797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T01:57:15.550335Z digest=sha256:4c4a96d0f057f6eeaae99f91651519263b7243388eb31516d2912a785364632b

Observation b53a9ff1-507c-4eca-b576-3096239f3fab · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.699283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:16eb828b1d40c2318095cc6510a0357bb0628300556a5bb88e50f39810ba319a

Observation e285619c-f356-468f-af29-8b43e4a227d8 · inbound

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment cites this paper.

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:11:31.576642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:11:31.576642Z digest=sha256:c831c9f4981dedbb6e0e21d579457bebd3c6344f9861bb5c1a2b35b5027dc5cd

Observation 78d4f218-73d2-4042-bb9a-9381a80e5282 · inbound

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation cites this paper.

Keyframe-Anchored Identity Preservation for Sequential-Action Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T16:30:00.510556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:30:00.510556Z digest=sha256:e4159bd6d09b24fabc26695e5e8b5a49822dcc32518c112751114ae6045aceb6

Observation c532743d-3314-4c6a-bf32-0df7f6622b84 · inbound

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement cites this paper.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:19.331303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:19.331303Z digest=sha256:1711cd1e23a66ea8c7839a42e2ac8213331836332d39d24fe582181901a9261b

Observation 36060856-f20e-4f5f-8a43-15ad641130d3 · inbound

Vera: Identity-Faithful Human Subject-to-Video Generation cites this paper.

Vera: Identity-Faithful Human Subject-to-Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:51.437202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:51.437202Z digest=sha256:ad53f18963d6a32bca0a1478f18306720c5679c166e1b8877f2ff3492f6bbfcc

Observation 893e6074-5bbd-470d-bdf8-5352e25aa5b9 · inbound

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling cites this paper.

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T13:56:37.838490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:56:37.838490Z digest=sha256:5fd1e6af58608fccaef012b061633a79991ec8feef48a32f5113f30a14de7210