Pith. sign in

Paper Citation Record · LEDGER

AnyI2V: Animating Any Conditional Image with Motion Control

As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2507.02857.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02857 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:25:29.212390Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T16:19:51.348169Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:38.518422Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy25
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1fed75c3-b20c-4f6b-96ac-78b1a717cd86 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

AnyI2V: Animating Any Conditional Image with Motion Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:24.234995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:24.234995Z digest=sha256:e5107ccab15676cf04ced3b5952a862f30df8f2a23dc4f12907dd206a56189af

Observation 88234fc6-27ef-4f33-83cd-dc35b5f64f2a · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Align your latents: High-resolution video synthesis with latent diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:34.460141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:24.323877Z digest=sha256:2944b21159754d20287d28546c61f410aae72ba86e5ae3b89eae09d1862c5a51

Observation efdf9c10-8737-47a9-a29a-4edda0d43b45 · outbound

This paper cites A unified 3d human motion synthesis model via conditional variational auto-encoder.

AnyI2V: Animating Any Conditional Image with Motion Control A unified 3d human motion synthesis model via conditional variational auto-encoder

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:34.328860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:24.413431Z digest=sha256:4f52de831e722b2379f8fe62f8a92a009d9c70b28ba994e49fb89975dd8cc4eb

Observation 47fb6f8e-8c13-427a-8c2c-c82e2e41c664 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:24.503120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:24.503120Z digest=sha256:f3b0b1fc1736c78893fea2808b26bd89ed857b854bb3cc33fc1802e0eddfc778

Observation 537051c3-c895-478e-8dab-e699d1e73474 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:34.189521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:24.594083Z digest=sha256:15b0fd63497f3253c438ad49163bc92bd1ee3d40fb356a55dd7459f0ec322598

Observation 78f7342e-fe75-436f-8721-6dfd29aad377 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

AnyI2V: Animating Any Conditional Image with Motion Control MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.997665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:24.679423Z digest=sha256:0a198003532c4b23f08b2a450fb1e196d23d7b976ecddb492a95a2a0c0597e62

Observation 8991ac9c-0e5e-49cd-86b0-9fe1b6ca30f8 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

AnyI2V: Animating Any Conditional Image with Motion Control AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:24.769350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:24.769350Z digest=sha256:bb3690d8abbf970e58e6d10dbff06dab653eecaa42eed3bd33c907b6beb64905

Observation 35b70089-05a0-4fa1-8c2b-22c36faedc0e · outbound

This paper cites Sparsectrl: Adding sparse controls to text-to-video diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Sparsectrl: Adding sparse controls to text-to-video diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.854495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:24.884766Z digest=sha256:bc31af222fb99b7c4196af5712b62bb4ac588d59874dcf3a3e7ee8a49925373e

Observation 84520573-1143-4c07-b930-d1d95e56baa1 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:24.979478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:24.979478Z digest=sha256:9566e3a31c8d63b5a6919359a2392dd2599c1fcb5e73400ed3469caafca11026

Observation 00467e71-d160-4147-9adc-40188cf446b9 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.057827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.057827Z digest=sha256:65afc1e472a0a04c1354c53dceacd1379ff97545ee33a7670ddea492ba8c868c

Observation c424eb1b-8aa8-43b4-960f-411cbbdaba4e · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

AnyI2V: Animating Any Conditional Image with Motion Control Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.140993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.140993Z digest=sha256:08c4a20e300396059b3ef89005befa1fc5fc6e88597b1d8d0d0b85122b12042b

Observation 4a3287ac-ad55-4777-a748-4f898526342c · outbound

This paper cites Video diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Video diffusion models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.684728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:25.240688Z digest=sha256:cae2b486ff774a5cfd09de33a6a4f3eef2d0855d57ab0991e72cdab69e737e6e

Observation c3b78684-d66f-476e-bafb-695ab7fa9ace · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AnyI2V: Animating Any Conditional Image with Motion Control LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.318105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.318105Z digest=sha256:a73a2a211fedbe7727615ddef54f9548fa4efc6e11b4cb328e7afe64b5f32c83

Observation 9fddd0ee-b922-46d1-9b73-51842255c65b · outbound

This paper cites Cocktail: Mixing multi-modality control for text-conditional image generation.

AnyI2V: Animating Any Conditional Image with Motion Control Cocktail: Mixing multi-modality control for text-conditional image generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.505365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:25.380684Z digest=sha256:16c92504c1530916352a8b5f82bb8a4dc9c7e690632f4432c9e46614ab4f524f

Observation 475fe6c3-411c-4d7c-be8a-a8000066e717 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

AnyI2V: Animating Any Conditional Image with Motion Control VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.466570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.466570Z digest=sha256:b94d0bbb6d2553b8de1bfc5dbb93e4ab597006c1339393238f322e299c567518

Observation bc66d003-8b0c-4769-bb91-666dc3eb6548 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

AnyI2V: Animating Any Conditional Image with Motion Control Arbitrary style transfer in real-time with adaptive instance normalization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.367870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:25.558706Z digest=sha256:369804ef73845bbdc1b1b2a3aacfe57dbed28d42c56dfaaf48e83ef42ddd6d92

Observation d19d4810-6fc5-4c9f-a904-61b8c942574f · outbound

This paper cites Cotracker: It is better to track together.

AnyI2V: Animating Any Conditional Image with Motion Control Cotracker: It is better to track together

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.236726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:25.679527Z digest=sha256:0ac08260248739aacd8060808ba066a7b7d2a6f9eee092d071ebe39d54603300

Observation 79574eae-be01-4584-ad52-ea507b4c5876 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

AnyI2V: Animating Any Conditional Image with Motion Control Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.815397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.815397Z digest=sha256:3c0cc9b6ed108c9063a9889eaf430f66a72b9ff50d9ac7d3ed3ddfddb5a06fd7

Observation b8490783-78ba-40ff-bd46-73591113df79 · outbound

This paper cites DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models.

AnyI2V: Animating Any Conditional Image with Motion Control DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:25.914025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:25.914025Z digest=sha256:4b8861b6fc1718185f00e99be98ae792a3e8f391365fff814a16f0bf22116c37

Observation f0a79e52-683e-471f-834b-a8b00b4bf615 · outbound

This paper cites Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis.

AnyI2V: Animating Any Conditional Image with Motion Control Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:25:29.851895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.042201Z digest=sha256:7bff9155919f0244fab6c1c15b7f917f1cd96e25cd2d4af1f17b9be970c89767

Observation d3b52798-77ba-4689-a296-1a3266b80d7b · outbound

This paper cites Image Conductor: Precision Control for Interactive Video Synthesis.

AnyI2V: Animating Any Conditional Image with Motion Control Image Conductor: Precision Control for Interactive Video Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:26.170396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:26.170396Z digest=sha256:7bc60a5c30436bb49bc61389b9f5daa88f502c310c451f6668db3bc8f3b71206

Observation 5feeba51-a878-426d-afe4-e1017b07d665 · outbound

This paper cites LOVECon: Text-driven Training-Free Long Video Editing with ControlNet.

AnyI2V: Animating Any Conditional Image with Motion Control LOVECon: Text-driven Training-Free Long Video Editing with ControlNet

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:25:29.716487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.239158Z digest=sha256:d1a06998e8020f62964405a090ab3870677cab35409c2157edd3ef679c430268

Observation 00524d34-72b6-4ca3-a573-6b5d1407a5dc · outbound

This paper cites Trailblazer: Trajectory control for diffusion-based video generation.

AnyI2V: Animating Any Conditional Image with Motion Control Trailblazer: Trajectory control for diffusion-based video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:33.012633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.320914Z digest=sha256:527f732679825cf233c9bad2022228db4c2ea2d6fd05cf7311ab783bbbe99087

Observation e796bd06-2fb9-498a-8b85-e3ccb432bbb7 · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

AnyI2V: Animating Any Conditional Image with Motion Control Some methods for classification and analysis of multivariate observations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.881148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.403459Z digest=sha256:7dfb92fbd955e74381a0f3b5b3c76e4e1dddd65a5990c9b52ec8264ccc8c14bc

Observation 661d2f45-91d9-4fc7-9bda-a2cced54d8c5 · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

AnyI2V: Animating Any Conditional Image with Motion Control Large-scale video panoptic segmentation in the wild: A benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.735514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.474151Z digest=sha256:d18b6249a0ae11faf51c03f3a6ed3963ade057a2f1c2c10131aec756d25cb601

Observation d73ee21d-b3cd-43af-97c1-0bea7ea1e6ca · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

AnyI2V: Animating Any Conditional Image with Motion Control Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.562331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.553258Z digest=sha256:33a90ce42f4a68b3854021eefbaa470cd1da5b7f472fbb683c171de8c5c7a85f

Observation 8526d6aa-bf06-4f6b-85a0-d2c9887d5ca2 · outbound

This paper cites SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:26.611739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:26.611739Z digest=sha256:ba89543f97098a339c8a48382113b66781bb5337e16a0719db61a7a5b8e8dddb

Observation 0afba912-7f85-4603-bfd0-d1ec67db126d · outbound

This paper cites Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model.

AnyI2V: Animating Any Conditional Image with Motion Control Mofa-video: Controllable image animation via generative motion field adaptions in frozen image-to-video diffusion model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.433431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.685839Z digest=sha256:b28b0333351ee668f7bd7b62a18485ecc8469f885574e9b34ebfdf8a4f407ec1

Observation 92e86803-561b-43a2-b383-9cdfb26ad7d8 · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

AnyI2V: Animating Any Conditional Image with Motion Control Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.218979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.760190Z digest=sha256:00520bba8e1b3b5deb915a7a63b3491f36745ce8c984dab7fadd1e48d260f14f

Observation 38b44cb9-bcc3-40a5-8436-4c7176e5d212 · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

AnyI2V: Animating Any Conditional Image with Motion Control UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:26.818684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:26.818684Z digest=sha256:d40a6e63ce8ac6d9e2e57a965e6c8b3682554bcb4a9a48aaa956b812b16969c2

Observation 42402254-5ab4-4997-be70-b199df2bf15c · outbound

This paper cites FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models.

AnyI2V: Animating Any Conditional Image with Motion Control FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:26.881733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:26.881733Z digest=sha256:92df1c4cbe68b88642b6dac0ad087913da713b03f2499010ba380357ad9e29af

Observation baed88d4-efb6-4686-a27d-15a66c5355a1 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control High-resolution image synthesis with latent diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:26.900124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:26.900124Z digest=sha256:465c89ea4abddef332e4d029bd3e179cfacda85a277d915f2fc334078bc720fd

Observation fd175d70-34e2-448a-b6f4-b2577bf1b0ce · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling.

AnyI2V: Animating Any Conditional Image with Motion Control Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:32.068727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:26.970264Z digest=sha256:f72972a9bba6acac22757add8b0bf470d83fdf50fb35fc3c20a662dd1495918f

Observation 15eb4f40-dd58-403c-b2c4-ed011779c61b · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

AnyI2V: Animating Any Conditional Image with Motion Control Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:31.963843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:27.078636Z digest=sha256:b80b99418ac3432ed7c482d8b4a2516827dc18b043943868457e2bbb44b29a26

Observation 40b743f7-bc70-444f-8b57-053c8fc6fc83 · outbound

This paper cites A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models.

AnyI2V: Animating Any Conditional Image with Motion Control A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.267209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.267209Z digest=sha256:6ef66af2ec233ddbd458c508e3da4abd5984ba0512a2053458f578479dc97b4b

Observation 32df7593-0501-4d21-bc58-60be5e71da60 · outbound

This paper cites Free-form motion control: A synthetic video generation dataset with controllable camera and object motions.

AnyI2V: Animating Any Conditional Image with Motion Control Free-form motion control: A synthetic video generation dataset with controllable camera and object motions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.408167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.408167Z digest=sha256:fdf945fd295125e410fd1ca31f2094738a743a5f100e7fdec3c9d628f4f4c6c0

Observation d3201c98-aac7-4366-8a95-f8f259a0d87d · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

AnyI2V: Animating Any Conditional Image with Motion Control Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.446167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.446167Z digest=sha256:5afb1f0a75817078dc76b8bdda81b716aefac5e86660e907124722831c2ca241

Observation b5292019-0b83-4721-be2b-ae7ff2d5fb28 · outbound

This paper cites Denoising Diffusion Implicit Models.

AnyI2V: Animating Any Conditional Image with Motion Control Denoising Diffusion Implicit Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.533479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.533479Z digest=sha256:f3a93647fea61a9f28feee7d44cbff2e87ad9eff4766cdf3b727c51ff361f86f

Observation f6a8d720-8439-483d-a44e-83d33bed5481 · outbound

This paper cites Anycontrol: create your artwork with versatile control on text-to-image generation.

AnyI2V: Animating Any Conditional Image with Motion Control Anycontrol: create your artwork with versatile control on text-to-image generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:31.798937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:27.623943Z digest=sha256:831dc2e8494f189881a0f311ed791a0e6b8e93c2cc39fc198601cdfdef473609

Observation 27c0b93b-146f-400e-91c5-0aa28818b3fd · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

AnyI2V: Animating Any Conditional Image with Motion Control Plug-and-play diffusion features for text-driven image-to-image translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:31.654118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:27.774197Z digest=sha256:10590532ab229d11a0c22baf81740e610df01c5a50ddc0c174ff0f727d98ff8f

Observation d3771b2c-ec8f-4c5b-af71-83e19304621e · outbound

This paper cites ModelScope Text-to-Video Technical Report.

AnyI2V: Animating Any Conditional Image with Motion Control ModelScope Text-to-Video Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.903864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.903864Z digest=sha256:153ed744c159d005b7b369e0025ca53038e62bb083de1b9325376dfc572d9421

Observation ef2b81e3-e487-4b92-a6da-fb15b3fc5765 · outbound

This paper cites Boximator: Generating Rich and Controllable Motions for Video Synthesis.

AnyI2V: Animating Any Conditional Image with Motion Control Boximator: Generating Rich and Controllable Motions for Video Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.032439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.032439Z digest=sha256:982de7b412c9604a159bbb612d865bb0306cd96db0c7d77bfebc5794a582a9f2

Observation d856d6b6-8afc-4eee-a71d-a74b60ec6519 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

AnyI2V: Animating Any Conditional Image with Motion Control Videocomposer: Compositional video synthesis with motion controllability

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:31.526763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.197918Z digest=sha256:65f2ad4510429f2927350fd049683158830af690d66eb94cf7871a088e7a3f14

Observation 37898861-1e15-4fec-b544-e391ee84143d · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Lavie: High-quality video generation with cascaded latent diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:31.177524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.282240Z digest=sha256:c9560f69d562209aae6a9dbe10382212fdd6e2167063e57ce18fa33161142532

Observation c9840de4-56e8-4fd6-9acd-8ff050e23d5c · outbound

This paper cites ObjCtrl-2.5D: Training-free Object Control with Camera Poses.

AnyI2V: Animating Any Conditional Image with Motion Control ObjCtrl-2.5D: Training-free Object Control with Camera Poses

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.341243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.341243Z digest=sha256:8664395484e6ac5a82e2382090871190b8f217c94cd6422603795d0ab1c49f4b

Observation 91b33beb-3f62-4ddb-969e-f0e49dd75f68 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

AnyI2V: Animating Any Conditional Image with Motion Control Motionctrl: A unified and flexible motion controller for video generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:30.904545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.415133Z digest=sha256:159dce1bc34fbb2b6dce701a7d49b9d1bcc68ef2378b47a19dd5bed37f8a79b9

Observation 396d7879-b606-43e1-a74c-6cba50ac1ae1 · outbound

This paper cites Principal component analysis.

AnyI2V: Animating Any Conditional Image with Motion Control Principal component analysis

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:30.631228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.525885Z digest=sha256:3247c6630f2f6682a6f7064ced1a5ff9c67b63ad58e672a75634af8739a6d2fa

Observation 39f036ab-5588-464d-8392-59af76ab3bf2 · outbound

This paper cites MotionBooth: Motion-Aware Customized Text-to-Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control MotionBooth: Motion-Aware Customized Text-to-Video Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.608683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.608683Z digest=sha256:c9724f77f6f3cc62f459798b2742b665ad5bf090cfc53a802f7734f601537755

Observation 15b79da1-3c97-477f-b11c-507de6af68ec · outbound

This paper cites Draganything: Motion control for anything using entity representation.

AnyI2V: Animating Any Conditional Image with Motion Control Draganything: Motion control for anything using entity representation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:30.383936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.647601Z digest=sha256:9a8a17b2c9e045f72a4684f97c7d88d2c50d053151d45da8a46b86391a261e29

Observation d8787f87-8def-4463-8b24-f76278b73f0c · outbound

This paper cites Video Diffusion Models are Training-free Motion Interpreter and Controller.

AnyI2V: Animating Any Conditional Image with Motion Control Video Diffusion Models are Training-free Motion Interpreter and Controller

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.728755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.728755Z digest=sha256:12ba6e861198c94999f2061cc49c6d766a20a403db67f29ecbf4c2979279c1d7

Observation 4ab67a9c-61ea-49c4-a657-7d1bc1f651b0 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

AnyI2V: Animating Any Conditional Image with Motion Control Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.790311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.790311Z digest=sha256:8e8a8bff5c70078185e8c2ef177f7d26a69274e65e54b6f82cbc444d5339d857

Observation ec32a8f0-e334-4054-882a-570f449f538f · outbound

This paper cites Direct-a-video: Customized video generation with user- directed camera movement and object motion.

AnyI2V: Animating Any Conditional Image with Motion Control Direct-a-video: Customized video generation with user- directed camera movement and object motion

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:25:30.116321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:28.847157Z digest=sha256:2de384303a8cd10c053666ae057258f5e8956fce6765ebf0be1f6da7cf2dee26

Observation 695822a8-cc50-409c-8cf6-d9736a99a273 · outbound

This paper cites DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory.

AnyI2V: Animating Any Conditional Image with Motion Control DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.946126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.946126Z digest=sha256:1f446c0721d79704cc9b5ed08e822caec0fec9c3fd6e6535193012e0a4080a79

Observation 5b748225-0f62-4468-8f00-57e762518ba8 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

AnyI2V: Animating Any Conditional Image with Motion Control Adding conditional control to text-to-image diffusion models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:28.992068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:28.992068Z digest=sha256:f448e46b39772dcd55d2d9da37db67ea8830e49f7bddce1e908dda46ab4e215b

Observation 1dc34cde-99a8-4f00-8daa-3face3f52316 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:29.057301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:29.057301Z digest=sha256:c62a5f703066403a4feeb679e8fa3b1ca59a3083e55c7e49ae6808aa2779a5f5

Observation 1351d6ee-1961-4baa-95d3-85db563283d7 · outbound

This paper cites Tora: Trajectory-oriented Diffusion Transformer for Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control Tora: Trajectory-oriented Diffusion Transformer for Video Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:29.127199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:29.127199Z digest=sha256:52ec4fda6d3a10337578e2d1ff47e8f4ac2da0ca236692a159c260afb7db1cb6

Observation cdcdc99d-47e6-4ea1-8476-9548690d5220 · outbound

This paper cites TrackGo: A Flexible and Efficient Method for Controllable Video Generation.

AnyI2V: Animating Any Conditional Image with Motion Control TrackGo: A Flexible and Efficient Method for Controllable Video Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:25:29.416173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T20:25:29.212390Z digest=sha256:28e08006af6bb44032bfd1dba2992e357a49cee68fa7a03f2368ab99a0fc7a9b

Pith citing papers

Observation bd2c33f3-ad0b-4654-8b84-00132bccf2f5 · inbound

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers cites this paper.

QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers AnyI2V: Animating Any Conditional Image with Motion Control

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:38.520077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-03T16:19:51.348169Z digest=sha256:a9a34c6b2199f032e2d9cd8927aeda1a2b18cbd0d3b0650ee6d6ae5188d872ad