Pith. sign in

Paper Citation Record · LEDGER

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 7 inbound Pith citation observations for arXiv:2501.01427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01427 v4

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:34:11.937640Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.028853Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:50:21.420628Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cda8aabf-5981-446b-b5eb-67244289d288 · outbound

This paper cites Video generation models as world simulators.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Video generation models as world simulators

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.743878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.568115Z digest=sha256:1e7c4ced3819a4136383ea079764a956c0552830d9aa462b65b99d9c8309b8ba

Observation 60660e59-0083-4dfa-ad40-2ad2a13cb8f5 · outbound

This paper cites an unresolved cited work.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:34:12.729979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.573560Z digest=sha256:ee25fbf91c8f66f098a501a526db9ceb8e84be1145f1660162d68e4762e4a1f0

Observation da1175b1-1e5f-4f81-b9cf-aad4ed7c6e9b · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.715405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.577667Z digest=sha256:55c33ddda4e9831ce4d82f3c409040265cc4800da3599f75ac6eace391bdcf1c

Observation b6e6ad9e-9fdf-47b2-b52a-755e591cee1f · outbound

This paper cites Anydoor: Zero-shot object-level im- age customization.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Anydoor: Zero-shot object-level im- age customization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.699867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.582016Z digest=sha256:2f7d535df7609fa53c6dd712eb57bca74d92fda9b339441aa4035b0a40e8ccca

Observation 0a8dd8dd-8b05-44a4-b011-55188da4e1ed · outbound

This paper cites Diffusion models beat gans on image synthesis.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Diffusion models beat gans on image synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.684985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.586408Z digest=sha256:4b7406b46e4d4ae8d60865cfa8df87988fd3defba391675319c4bbfb6719b1aa

Observation 195201ee-0812-4232-9dea-d777359f4dcc · outbound

This paper cites MOSE: A new dataset for video object segmentation in complex scenes.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control MOSE: A new dataset for video object segmentation in complex scenes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.666628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.591184Z digest=sha256:6c90b1fd82997fa6d446b950abee25cb2967f1e0e6140e050f44c6a1fe892e9b

Observation 9edd5551-1e7f-4aaf-a0b6-e05f06ff2304 · outbound

This paper cites UVO Challenge on Video-based Open-World Segmentation 2021: 1st Place Solution.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control UVO Challenge on Video-based Open-World Segmentation 2021: 1st Place Solution

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:34:12.257287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.595100Z digest=sha256:61794fb83fbab270bf9c58855dfe7cb74471599a4300b9ecda3e7ce71dcfd3b8

Observation 3bca78bf-d91a-48cb-be93-7878b435623e · outbound

This paper cites ViViD: Video Virtual Try-on using Diffusion Models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control ViViD: Video Virtual Try-on using Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.601738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.601738Z digest=sha256:c50865ce61bd7649713f93975452e1db0a5b349ebea13546f60042da1451d2e5

Observation da8eba32-caca-4e45-af01-1b099052641a · outbound

This paper cites Tokenflow: Consistent diffusion features for consistent video editing.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Tokenflow: Consistent diffusion features for consistent video editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.651448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.606722Z digest=sha256:d386d96819c3674e701b3918e69ed982b741f98771639c29b9208b48c5f96538

Observation ab1ee9c5-4983-450b-b84b-6e5c9522e0a6 · outbound

This paper cites Videoswap: Customized video subject swapping with interactive semantic point cor- respondence.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Videoswap: Customized video subject swapping with interactive semantic point cor- respondence

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.636531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.610714Z digest=sha256:c73251bba8ea21e8449332bc87197b3de850fa1a3aafb968bdcffcc0224e41f8

Observation b6f112a3-5fa8-4cdf-84cf-59f9f63152a3 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.615433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.615433Z digest=sha256:b321e8766dfb9fb5d1a0ee0dc36f80d2fbd8664fd6d0e896388835cb34a45abd

Observation 7ba576b8-7bea-4de3-809a-bfbedd9d2954 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.620489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.620489Z digest=sha256:7c237eb65b056e27e69e100aac88730d3714cd69cb5b49bb4fe912b0c613a58a

Observation 19e9a463-44ee-48c9-ba3f-55f9fa3f7353 · outbound

This paper cites Ground-a-video: Zero- shot grounded video editing using text-to-image diffusion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Ground-a-video: Zero- shot grounded video editing using text-to-image diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.613287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.625653Z digest=sha256:4c986f28ca8ec628b228cb96ec64fa5551f957d0e7d2f1a58e9c0b09e9e738ea

Observation 68fc0ddb-3724-412b-8010-b1776f68558c · outbound

This paper cites DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.633226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.633226Z digest=sha256:f8442724f5ed9f40d40a34997309c4cb8c4b33463611bdf07f1786bbff70fc15

Observation 3df1152d-043b-429b-a110-fb9d1c04e939 · outbound

This paper cites CoTracker: It is Better to Track Together.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control CoTracker: It is Better to Track Together

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.638335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.638335Z digest=sha256:3d6d6185a97836e1f3d7f76ac979924108be9d9db40018ae2d5315e22869ee23

Observation 7c2b03ef-21dc-4636-9e2f-55ada3f6166b · outbound

This paper cites Segment Anything.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Segment Anything

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.642702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.642702Z digest=sha256:93ae74a976a92125b81f2946e681d340f5173e8a518503403f979d3a5e83e7ec

Observation 99e8d3e2-c664-4c39-9449-38233b4976fc · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.647567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.647567Z digest=sha256:0c9634b98bb9a315fe5ae0e63ff92b8c41b4d47117d58dd78c4900c2236ceaf5

Observation 964e80d0-cafa-4f91-9fa8-83d1832451e3 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Multi-concept customization of text-to-image diffusion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.599109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.652483Z digest=sha256:b0ec467162c8ec6cfd48559dcb686ca5c75a5e8d5f42939e895be497df94055f

Observation e0c78b4c-c02e-417f-8093-3bf78c38fe30 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.656783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.656783Z digest=sha256:542e458c3279fde12ade3abefb7e756608bfca42fc730acceea186a22ade05f4

Observation 264b61ef-b3e0-4bbc-81b4-74f5079fa6ad · outbound

This paper cites StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.661933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.661933Z digest=sha256:9df07f185c093eefae67267fe9972c26d9b732decb7f18780c7ce430ebb7c549

Observation 38da6899-f887-460e-ae70-f16fe26bff85 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Video-p2p: Video editing with cross-attention control

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.577045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.666554Z digest=sha256:94e40f7a61464fdc76528586375b03275500e42ebe6fdea03bdc89c799182690

Observation eff90762-4f0e-48d1-92b2-c0ff321a49ae · outbound

This paper cites Large-scale video panoptic segmen- tation in the wild: A benchmark.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Large-scale video panoptic segmen- tation in the wild: A benchmark

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.560220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.671448Z digest=sha256:0d76310c86007bfb6da4fcb0c98cf125bbd28d57456f5f0bef62b59242e13ece

Observation eaacfcc7-079d-4408-9f0a-bd957747b3ce · outbound

This paper cites Vspw: A large-scale dataset for video scene parsing in the wild.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Vspw: A large-scale dataset for video scene parsing in the wild

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.546705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.676161Z digest=sha256:d6f863cf9b6ebcc9fc4b577671f98a230c4ce10c95118081594d3f610e3b4115

Observation b4ecb72a-b4b1-43aa-9d83-3fe779e4d95c · outbound

This paper cites ReVideo: Remake a Video with Motion and Content Control.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control ReVideo: Remake a Video with Motion and Content Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.680264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.680264Z digest=sha256:e3bd459cc1c05ce457ccee5e9fe938ea458887bd21d0259298332999df876f72

Observation 39cb9cca-f11b-4e97-b225-9d89c61640f7 · outbound

This paper cites an unresolved cited work.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:34:12.534248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.683947Z digest=sha256:1c7035efc5db96727f2b8079b0d8628308bc7009381a31db7d24dea85844cacf

Observation 1a108ee5-d1b6-4be6-8121-a609b16ee816 · outbound

This paper cites Codef: Content deformation fields for tempo- rally consistent video processing.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Codef: Content deformation fields for tempo- rally consistent video processing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.520030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.688467Z digest=sha256:551c179c73ba79513080ae61104b46e0e4d64b495d32e89b0d5a0797da1e0dac

Observation 3b1d3ff6-4f27-4c1e-8a44-5d334d24f5d9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Learning transferable visual models from natural language supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.505987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.691689Z digest=sha256:503dbff65e2e614504ec5e45e49cd12a6fd1a7c29a9987f66ad780c320843fde

Observation af34985c-b868-466c-b4d5-c35007b2001b · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control SAM 2: Segment Anything in Images and Videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.696158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.696158Z digest=sha256:9fef8672766f4a12d62e4b91c47ba3c5ced11e1c6a6d170edb6677beea22a880

Observation 2c22a101-9b4d-44ac-ad53-f057aa8e1c90 · outbound

This paper cites Consisti2v: Enhanc- ing visual consistency for image-to-video generation.TMLR,.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Consisti2v: Enhanc- ing visual consistency for image-to-video generation.TMLR,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.493034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.701219Z digest=sha256:1badd33042b81aacab3bb4c1efcbc1cf5f750f5248b7c4acaa08f6b3df270ab5

Observation 51df3bdd-3716-4c22-be50-3cdefa6501c1 · outbound

This paper cites Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.480550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.705116Z digest=sha256:478f0821d4e3d96c811f9e66c2b216b9063053bb6161ba25df39444281e7867a

Observation db0f8dd2-4523-4f11-a96b-627c0f5ecb72 · outbound

This paper cites High- fidelity guided image synthesis with latent diffusion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control High- fidelity guided image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.468441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.709145Z digest=sha256:a0ea5ef0ae699708716cd7226055ee7dfaacc6fa87bf4c2f7660d085161c79b9

Observation 7997fa17-c2f1-44e8-84af-d6fbe33bfb07 · outbound

This paper cites ObjectStitch: Generative Object Compositing.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control ObjectStitch: Generative Object Compositing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.712961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.712961Z digest=sha256:e9e04a256cb37d390e8d91fd2e0e62c50590934376c083544c9f83e40bcda038

Observation f8385fef-a939-4598-8794-803d2870c701 · outbound

This paper cites Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, and Daniel G.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, and Daniel G

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.454355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.716886Z digest=sha256:baa6eef7ee5bc4941cef8d69a732bea60817af8c2a316ddf5075d09452b2464a

Observation fa687aa1-769a-4813-a153-3bbdd6d3382a · outbound

This paper cites Weighted-to-spherically- uniform quality evaluation for omnidirectional video.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Weighted-to-spherically- uniform quality evaluation for omnidirectional video

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.436068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.721090Z digest=sha256:dd03ce2cfba8312b720f200786c4788afc7bba87df7aa1700f273f1c8ff0bd33

Observation eb54f76d-9556-49e9-9c0c-13281cbe3674 · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.724998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.724998Z digest=sha256:da681da0ac4c6e0e6e9d9ccdfc5c1e78773fbdc91accdf68af04a1e5a54e161e

Observation e1ef8393-4dc1-49ec-8980-60702d488b6c · outbound

This paper cites Motioneditor: Editing video motion via content-aware diffusion.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Motioneditor: Editing video motion via content-aware diffusion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.409157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.729248Z digest=sha256:25e1a8080de54e86281a1e985af602ce19af9d5378b6aebf1f59b99b4a299275

Observation 396e9e9b-3a7a-49ea-8271-aaf8a2dd8dda · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Videocomposer: Compositional video synthesis with motion controllability

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.395117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.733196Z digest=sha256:def90c6338287a1d00c0fa55c820c892703e66b67207e1f5a8f5048dfa6abcf5

Observation d33022ad-a9c2-4b5f-b049-cc7f08d1a894 · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Dreamvideo: Composing your dream videos with customized subject and motion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.381093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.737129Z digest=sha256:c8ca60b207438b5110c4bc8e6b27c7d66ece0627a214ff86844b96d65e10e1ac

Observation 61f5e7a4-92cc-4eee-9b10-3c08f06338af · outbound

This paper cites DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.740984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.740984Z digest=sha256:403d0e0ce1ebd8e7bf8cf99714584785675241307ac974034d0c83b4735acba3

Observation dd6ad1b0-48fe-4d6a-8d7b-509c591472b0 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.745075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.745075Z digest=sha256:dff2eda9a5cfa959215313a0158f551a3c3da19520e05fbb1ae2a6a8afbdc28d

Observation f041ef84-ea84-4496-8842-857360f9fe29 · outbound

This paper cites CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.748709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.748709Z digest=sha256:a5150e615249a6ddd6b041a274239d447951c79cf60ff941e1ff74e07e593037

Observation 1a2f0180-c1ce-4388-9f1e-56f49ebc10c3 · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control I2v-adapter: A general image-to-video adapter for diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.358879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.752779Z digest=sha256:a9f099b89953045c7897c20b949b592a962fc086299b7e0bfeed156ec3bb2c95

Observation 51d3c264-fe00-44e2-9fa7-d301016e16c5 · outbound

This paper cites Paint by Example: Exemplar-based Image Editing with Diffusion Models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Paint by Example: Exemplar-based Image Editing with Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.756265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.756265Z digest=sha256:2c38be8b74cf8acc5fece2c0256e3ffd1b8651d2181e41b1feb28b277da399e8

Observation 40fd7eb1-d4a7-4065-893e-cc1db3d1a2f8 · outbound

This paper cites X- pose: Detection any keypoints.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control X- pose: Detection any keypoints

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.345995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.886620Z digest=sha256:d257d76f26d35ce4c7c9560a291f3f97ac1d68d77856a1fa1eb7bafcce0e043c

Observation 6243cc6b-0349-45c9-a6b9-bb72d82b61c7 · outbound

This paper cites Video instance seg- mentation.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Video instance seg- mentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.332381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.891368Z digest=sha256:e636bc6c7997874e28f8d16873ea9ae21e37c7af900d3f1e27eadeb051f39c28

Observation ff36d78d-270d-4ff1-9d26-a5f638f69c2f · outbound

This paper cites EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.895152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.895152Z digest=sha256:33659a72b71d6600dfee6a6609a8a1f08bcfcfbd6d2d425983f300411fa33065

Observation 3409f0f5-5bc7-457b-a38c-385257913928 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.899864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.899864Z digest=sha256:d43299212a90e59c62e87ad6d867772eb19428cdf0714227c0bcc2e7301d8e04

Observation 5e117ba5-4c15-453a-a874-4bbfe2291ac7 · outbound

This paper cites Mvimgnet: A large-scale dataset of multi-view images.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Mvimgnet: A large-scale dataset of multi-view images

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.317263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.904144Z digest=sha256:462cb71c5444a9d59248fb94d4bb13fcab9114f93c8beed3418d6984234c43ba

Observation 9d16e947-c342-4165-aba8-618234d4dc4f · outbound

This paper cites Customnet: Object cus- tomization with variable-viewpoints in text-to-image diffu- sion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Customnet: Object cus- tomization with variable-viewpoints in text-to-image diffu- sion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.303410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.908160Z digest=sha256:09a8957b40544ad7e734182233e3d3bca179404a04d68445fc35be55cd82d51b

Observation ca9f5c05-06d8-4a4f-95fb-c0c02fbe9f64 · outbound

This paper cites ControlCom: Controllable Image Composition using Diffusion Model.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control ControlCom: Controllable Image Composition using Diffusion Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.913066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.913066Z digest=sha256:5557ea8b193fcf553693ca986c4dd75203514620bf32b5f2b874d7a28798ef89

Observation c13494a9-88a4-40fb-89a1-ee3e923516e3 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Adding conditional control to text-to-image diffusion models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.923470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.923470Z digest=sha256:4e1e8cbe21afef2c9f68a61791934b8140fefe467290ad991fdcdb497908feec

Observation c59ab23e-f530-4770-80a6-d8227d1199c8 · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.928216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.928216Z digest=sha256:f50fe17db04c18684b06a92e1a4ce2b484e4bd958c4ac1085b1b841a830a1310

Observation 19896ce2-b18b-4313-b8e2-5ffd239cd0f4 · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:34:12.280594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T22:34:11.932824Z digest=sha256:03a4736a690f410bd5b36db9544553ab39cf2b779b3cde4ff8b4a1107212bbd7

Observation 8cfe883d-fa0b-4cb2-ba43-0a8f01038dc1 · outbound

This paper cites CelebV- HQ: A large-scale video facial attributes dataset.

VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control CelebV- HQ: A large-scale video facial attributes dataset

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:11.937640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:11.937640Z digest=sha256:8752cdece825993d0a195bb79362c9872172e466f3bb46fe085bb6bb3c4306d4

Pith citing papers

Observation 34f7f8fc-495d-4c0e-811b-700ac56c281a · inbound

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation cites this paper.

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.028853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.028853Z digest=sha256:b149bfeee14245c4f08955cfa59c318f9b569503bd460291205848e7492abf5d

Observation a4a43647-10bc-4046-bb08-15728a47768a · inbound

UNIC: Unified In-Context Video Editing cites this paper.

UNIC: Unified In-Context Video Editing VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:42.902061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:42.902061Z digest=sha256:508ffdead6cd1f4e0b4b31b5e8a30187527e029263202805f7ebfbb13bc834f1

Observation 09d6f6fb-0c56-41b0-aef6-97783641af11 · inbound

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing cites this paper.

O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:35.467169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:35.467169Z digest=sha256:d799e74f1536cec9d7fbe28f6ce63f87f565cca15ec9a00a4e06604642f84196

Observation 67791fdc-e535-4b24-872b-76b2d53c339e · inbound

Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer cites this paper.

Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.017328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T18:36:03.018873Z digest=sha256:b89d6105f29a9820f95b50c3bfc2a502265bf753801e895feebca72d09caa8aa

Observation 3e82e298-36cc-4a5e-a4e4-dc517c089e21 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.586351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:10f1d0fa5bb4291c217a0f00e14d345873fb72b7f9cb110f6979326080a8b5eb

Observation d545aff9-9179-4a9a-936d-54882d989ab8 · inbound

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing cites this paper.

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:01:31.496416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T15:44:25.021995Z digest=sha256:bce71a9fd17cffbf748233141674491282c77f6c540802f99f2861211e9f615c

Observation 6851de04-5895-498d-a226-4861879ea9cd · inbound

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion cites this paper.

SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion VideoAnydoor: High-fidelity Video Object Insertion with Precise Motion Control

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.424085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T04:46:50.749223Z digest=sha256:596999f2a58b664badc5e2884fbd2a49abaa3a0fe7da5912a422eb050ac4f2d2