Pith. sign in

Paper Citation Record · LEDGER

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

As of 16 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2501.03059.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03059 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:01:57.121114Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:17.425943Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:05:51.755490Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a79eec42-9284-4ab7-b538-b06d8183c268 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.903581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.903581Z digest=sha256:4c408d6b16990875e369d7396afadc178d9713e58251a72b023fb0be7428a616

Observation 53b577ea-56c4-4119-b4a2-dea4fddab751 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.908137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.908137Z digest=sha256:c86c1802e6f367b4d4697b907d6179ee3e642120d634a09d5bf24d4dad520a9f

Observation 026bdaca-e3e6-419f-bedd-d90ba08bcb12 · outbound

This paper cites Spatext: Spatio-textual representation for con- trollable image generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Spatext: Spatio-textual representation for con- trollable image generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.762151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.912145Z digest=sha256:fad784f106eed6309d960a243da8047f30178c0e3b1bc07440b7c738321f8a56

Observation 6545e6c9-05f6-45cc-b5ff-d5e1d6aec3b7 · outbound

This paper cites Improving image generation with better captions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.916116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.916116Z digest=sha256:386c600d5726c07591d0664ec8f5329c49321559415cab9524186eed9601dd3f

Observation 1e61b7d5-daca-40de-a087-2c92a7c4e001 · outbound

This paper cites Understanding object dynamics for in- teractive image-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Understanding object dynamics for in- teractive image-to-video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.739754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.919447Z digest=sha256:91da17699b012386b5b65bca1918e79f76b5da44e200a37c33d593921439f7fa

Observation 3068459e-e59b-4a28-883a-7bf4ccd5a9f3 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.922872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.922872Z digest=sha256:2e5b7512abe566ba99ab501c7e63c50cfcb945cea9cd004cf7ec52290ed7d744

Observation bb540953-09a0-4a6d-9f82-963835373585 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.926465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.926465Z digest=sha256:f1b9d851f17c8ed0d3e12657a99aa622171fc311adbb4df72f3ecfd3132de805

Observation b4d8372c-59c9-43c0-89a9-54498e30b8bd · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.710141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.930307Z digest=sha256:36219152d4dd380fa217f5d4d0ccf56c56900bfe952b919c648903447d494b40

Observation d176fb41-f0cb-4dae-bb5b-44aed1de8e5c · outbound

This paper cites Videocrafter1: Open diffusion models for high-quality video generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter1: Open diffusion models for high-quality video generation, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.696417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.933896Z digest=sha256:ce6f1b988f2691fde76a69cc6eceb2760dd9fc65374a8f24f91bf816601206ce

Observation d4752850-775e-48dd-b57c-61b559bcec21 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.683744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.937556Z digest=sha256:81521d64473dedf3df667777ca434bc7d75e3f0c94e19af1dd70238f9a00866b

Observation 59486bbc-3e52-4d5f-8ad1-798f0469ce3d · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.941099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.941099Z digest=sha256:69034da4d87827029f81f9a861a46cca73f14cfe6c2d88e68dc2ecf529ad58fb

Observation 241d01f1-fc55-4112-b15e-f50fd3eb7c06 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animateanything: Fine- grained open domain image animation with motion guid- ance, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.667603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.944984Z digest=sha256:4ccd87f43f643830cdd70ac16098e19699fb815b249102f767e681dd50f3bb31

Observation 339b29cf-d819-4082-a1ad-ba821192fe25 · outbound

This paper cites The llama 3 herd of models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The llama 3 herd of models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.651638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.948143Z digest=sha256:82c2fd1130a07628744da28573646018011b396bf491a1f38eb9b060c697f423

Observation cf8d510e-3057-4111-bc9d-7b91f7cf6eda · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.641005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.951413Z digest=sha256:e6eef1b7af0ece6cacf489df4bd3829a0bb6a45e912c56fe14ce9b53f7ac100e

Observation 4e319397-44a7-4626-88ff-29f3d1e94af5 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Preserve your own correlation: A noise prior for video diffusion models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.954856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.954856Z digest=sha256:c5a9af7cb163e0bc204099317bc0a684f9a22389189ad727965193068daa1686

Observation f1465803-95ba-4758-900e-5490b1021b54 · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.958288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.958288Z digest=sha256:6db05192bf65079f2daa7620004788339f4beaaf3cab8208361a7b0a74d0eb27

Observation 6f791340-10ec-4842-ab9f-a80fa5685c64 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Animatediff: Animate your personalized text-to- image diffusion models without specific tuning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.623359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.961966Z digest=sha256:2fb4548cd67ec553128df6d0209c90970ca7a79383167a1627b80f26c2f6b377

Observation 794d29f6-db4c-4e8d-b549-c34e33193d3d · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.965670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.965670Z digest=sha256:5117b41cdd99ba36140e3cdfaa6a6b0350bb147f37c7d7cb4e8dd06e6b47797f

Observation 59eca75f-cf40-44da-8858-40a590a90645 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.969711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.969711Z digest=sha256:ad5a9ba530472ac3979d8306151e18a24ada4f684823d7a6ece15ac2af61e141

Observation 43a9cf77-7e3b-4d9c-ae18-32865dfe6879 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.973238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.973238Z digest=sha256:deee971952610b70134feae48f554a42040d21f6e7439fa02e22e2d24b912e89

Observation 35d94a97-3fac-43a3-b083-25da50e759fe · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.976611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.976611Z digest=sha256:3e69d30d3c9590d5386290d8be8f6e11d472c76d30b7ffc2f81712f9be2959b5

Observation 6bb44c0d-11ac-497a-b775-f095f605d541 · outbound

This paper cites Video dif- fusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video dif- fusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.980424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.980424Z digest=sha256:6c83ab0b66378d8352c9e98ca34e288fa5b9ddd57e356ca27081b168fe7c4b73

Observation 8bf0d250-e71a-4f02-91e9-a6dde78bcb7c · outbound

This paper cites Auto-Encoding Variational Bayes.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Auto-Encoding Variational Bayes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.983661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.983661Z digest=sha256:7da692f379a7bfa31c84db09cf056a70eb4074177772f500262e417c3068a54b

Observation 1e2dcf34-dc34-409c-9fd3-2beb71745ce1 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.986615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.986615Z digest=sha256:49989b6dcc24d558b432ef15c494c7b6fae1291465c814ddfd132f7bb92c9f90

Observation fd5142c5-1752-4337-8487-7e606761e9d2 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Gligen: Open-set grounded text-to-image generation, 2023

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.599310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.989461Z digest=sha256:9448593bebd919788c0ea23d7c1e971768e2e19cc1614d7bdff4cbb69e36d241

Observation ba225d87-1e3a-4e6f-bf43-f36c23d0f811 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Common diffusion noise schedules and sample steps are flawed, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.587743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:56.992375Z digest=sha256:d970b74a857ba938fbb688e3483f83ea327b2149570ea81407c90d7e335798c5

Observation b4a4f9be-6999-4bef-8900-5f709a9707fd · outbound

This paper cites an unresolved cited work.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.995450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.995450Z digest=sha256:f4dda4d193b5bc4922fb60b9a8755e42f1ec3781a79075fbd7f0dd700a4c0a57

Observation 465be81b-491b-4b76-a33d-52988a93c56e · outbound

This paper cites Rectified Flow: A Marginal Preserving Approach to Optimal Transport.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Rectified Flow: A Marginal Preserving Approach to Optimal Transport

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:56.998298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:56.998298Z digest=sha256:b68dc3b563ad2debb2f67ef1dc6ff3d75c29e4e797d6c89d6a9b1acb61370fe1

Observation 26e1a02e-49a8-4d0f-a258-a86e57f40f0b · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.001678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.001678Z digest=sha256:01c4c20d8ab81052e2f64585b79b72e49e77ec6c5d0293b9d407150a895bb236

Observation 1c944c83-5a2c-4b2b-ba12-af961d62b9c7 · outbound

This paper cites Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cinemo: Consis- tent and controllable image animation with motion diffusion models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.563305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.004917Z digest=sha256:670c41d79b109d8d606d18553408dc0a1d250b878848560de2ab0b91462892cc

Observation 2b2397d9-42a8-460d-a8d7-1570d3ddb895 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.009271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.009271Z digest=sha256:4450925ceed608f0b1772d83675c38208ec06edbbc9f68629f7579b9995582eb

Observation 2e2167d8-cb15-4b8d-954a-04158ab401ab · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.552663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.013032Z digest=sha256:685146171ff5fa0be71a8dcbe5122436bd670f93cc4e5b2df6571f15230d1f33

Observation 25e0d8ad-9146-4f52-8e0b-26e12dd0aa7d · outbound

This paper cites Compositional text-to-image gen- eration with dense blob representations, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Compositional text-to-image gen- eration with dense blob representations, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.542735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.016249Z digest=sha256:2f7211234f1d2faa36b163e4c72c1c93e04d1d1e9edbca1f4ff53642d14e7028

Observation bbdd92b7-fdd3-4aad-b416-20b7ce96c201 · outbound

This paper cites Video generation models as world simula- tors.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation models as world simula- tors

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.532320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.019827Z digest=sha256:9d3863da9e6d6c2a544ef067383ca21f4a2f7836cc45ba4564549150eebea7fd

Observation 89497550-e4ca-4777-808e-45c596c84c8b · outbound

This paper cites Video generation from sin- gle semantic label map.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Video generation from sin- gle semantic label map

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.521339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.023517Z digest=sha256:b82d13855b83fb658675e9111342ff55951fd240b751f01ec31f5398b7252f9c

Observation 87207992-9770-43ec-9487-d23ab702bce9 · outbound

This paper cites Scalable diffusion models with transformers.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Scalable diffusion models with transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.026576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.026576Z digest=sha256:6368fcc6c5f89e44b9722af5758a24726b4f83f6413f8e0997e292c11ec726e4

Observation 4a42d2a5-590c-46f4-beeb-241a526e16e5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.030059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.030059Z digest=sha256:fb01529079fa1367fe6a4c90d55cce2f2d4c367e0ba530753d31f492513268e2

Observation 7dd62e78-f0a7-4dc4-9b9e-005dbf124fe3 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.503944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.033703Z digest=sha256:b63e96fc1aa586d580dc838c13c76c3fadd46ec18d20e5e2a9315426b2952fd2

Observation f453eeed-d4f3-4658-be80-4da6ce43f96f · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Learning transferable visual models from natural language supervision, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.036908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.036908Z digest=sha256:51ed18250259b057536235fd29d2c63e246445c2762c6b49b669c49d4edfb5f5

Observation 20d15cec-fdb4-4863-997d-6ba41dcf599b · outbound

This paper cites Sam 2: Segment anything in images and videos,.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Sam 2: Segment anything in images and videos,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.040095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.040095Z digest=sha256:6bd30929f61cf050d6151e09db9de9c184a9b30e52c2fda955b60e5dd1001efa

Observation 6a0e9035-f195-4696-ba21-d4e43d5c5ee0 · outbound

This paper cites Consisti2v: Enhancing visual consistency for image-to-video generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Consisti2v: Enhancing visual consistency for image-to-video generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.480654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.043751Z digest=sha256:a732a0c6333279ff0360bdfb3b1971706174c673187726f9e6eac8c82636d9ab

Observation e05b1915-93e3-4c75-9963-b40e97d5e580 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.470475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.047344Z digest=sha256:e40307195abf9830b43c41643320250b5dad1939dc588002a3689ad3c7d7370a

Observation 660f2410-fdf5-4973-b83a-6a28e00e162a · outbound

This paper cites Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.460028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.051467Z digest=sha256:f0da94c238b3f9545c5fa25024026627a2e94bead11c36af952e0c3561642684

Observation 7c554da8-0fcd-4a0f-b428-c38a9cd680d0 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Make-a-video: Text-to-video generation without text-video data, 2022

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.450239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.054893Z digest=sha256:ac7ff97e42aa06611b5bb9fff96c47fd8c207d531af299815e22de982d0e3216

Observation 556750ab-c2a3-4f8c-91aa-0055e0b78c60 · outbound

This paper cites Denoising Diffusion Implicit Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Denoising Diffusion Implicit Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.058078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.058078Z digest=sha256:4b11a21b21111bba8b71b51f0a0df0202877820e45a15d526ea5f5959bb8f4c0

Observation f0bf9f6c-004e-4650-9f91-47a4fd3ef298 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.061614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.061614Z digest=sha256:5b691d062b5be35d6fa5721f3700a10e25978265d9cd4838177271627c4717b0

Observation fb07f7b0-7508-4810-994a-3049dd0e15da · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow, 2020.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Raft: Recurrent all-pairs field transforms for optical flow, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.439948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.065341Z digest=sha256:816057579f80cd300e23587b445c1013da248598136501ced4e591216690d37e

Observation 979d90b6-12b3-4aa0-b422-001609b75440 · outbound

This paper cites To- wards accurate generative models of video: A new metric & challenges, 2019.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation To- wards accurate generative models of video: A new metric & challenges, 2019

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.068199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.068199Z digest=sha256:14d8727cb1707ecbe9472c97ed7f176eea0eb4e5e9ccc22740a87962af06a348

Observation 152c2249-a9ce-4d79-8163-89c3773a278d · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation ModelScope Text-to-Video Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.071166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.071166Z digest=sha256:b18c87c94fe47d709a74980374ddb4b1974743db5788c51a4a790f7313bb98f0

Observation 0060c327-9d8f-41d8-adbe-7add5337125d · outbound

This paper cites Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.074328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.074328Z digest=sha256:d8ff53c6ebd73ee54c93509136cf436a4ea63467a47588d34dcd0ef26b2fff1d

Observation e504a121-af29-4c76-9b00-cb055b0f80da · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.078986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.078986Z digest=sha256:f232f6541648e3e1fdbae805567a9b48ec1a72ca0a20fcd92f842acc2be609a4

Observation c74e4b38-fd33-4197-b3c8-0f0724d36a24 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Internvid: A large-scale video-text dataset for multimodal understanding and generation, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.082626Z digest=sha256:98ccceaf84c349558bba7bc20630e03db378c39dc9171207e89eee2293a5db93

Observation a6138c38-078e-4e60-9570-b216a2aeeaea · outbound

This paper cites Cvpr 2023 text guided video editing competition, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Cvpr 2023 text guided video editing competition, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.086032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.086032Z digest=sha256:2f4e194ac2684b8be6401afa1d17578efd3847da1053f0f4def8599db443f861

Observation 3e80e4cc-1dec-4c74-a97f-0de74b0da988 · outbound

This paper cites Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Dynamicrafter: Animating open-domain im- ages with video diffusion priors, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.407001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.089488Z digest=sha256:776893ffe0ae1044b0e6e5a80c82a788aba0cad3a3aba308a89b285529a6446e

Observation ee53ac5c-e98f-45e2-81dc-a4a741fb77d6 · outbound

This paper cites I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation I2vgen-xl: High-quality image-to-video synthe- sis via cascaded diffusion models, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.396106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.092834Z digest=sha256:25b1ebf5c6075c2196de2b09cae480ad42981e3f57609236f4642a4a6d4842cd

Observation f2cdff7d-03cc-484b-98ba-aca83a4d4f89 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:01:57.096246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:01:57.096246Z digest=sha256:e98e8181a82fc0d6934ec28289aa5f6a1b556928f4cca62c2cb67442e2463dbf

Observation 45863543-3017-4676-844a-782aeaa0ff15 · outbound

This paper cites Qualitative Comparison of Masked Attention Mechanism Fig.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation Qualitative Comparison of Masked Attention Mechanism Fig

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.386219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.100257Z digest=sha256:bfe66d03446183f113d82a89ae68dcb58d0d4cdf245365a7a3bc4c7b76f944c2

Observation 2c53be84-e420-4376-9e28-fbfd94990343 · outbound

This paper cites 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, our pre-processing pipeline ex- tracts a motion-specific prompt, cmotion, from the input text c, using a pre-trained LLM

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.374870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.103605Z digest=sha256:2d20ab52d0228acde45ef06b55608d3250599bfd7f9f4d27601e45e66ca1b58d

Observation a1db1218-8197-43f9-ba37-0f12a174fafb · outbound

This paper cites 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation 3.1, the pre-processing process be- gins with extracting motion-capable object prompts from the global prompt c

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.362362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.107151Z digest=sha256:27b15339459ecaef9b84ff950951f713f1a7dc3cfa1d83ec37bb86d24e7fc128

Observation 4ca3f9d3-c478-45b9-890a-0d95840fbab9 · outbound

This paper cites First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40].

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation First, the initial segmenta- tion s(0) is extracted from x(0) using SAM2 [40]

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.351616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.110565Z digest=sha256:9359c12aba02e98bf3667dc2ab4f5ae50d5b376de0855264f3c37e449cdfcfcd

Observation 9ed908a2-eb82-4c35-8cdd-cba69f8a898a · outbound

This paper cites The first is the U-Net architecture.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The first is the U-Net architecture

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.341696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.113763Z digest=sha256:492e60f4ed934570e61dc81989ff79cada1e96778695057063cadd3135ace7db

Observation b01e8fac-5167-4c7d-9b03-3d4a8a2b7c53 · outbound

This paper cites The filtering of 128 videos, out of the full SA-V dataset, involved several steps.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation The filtering of 128 videos, out of the full SA-V dataset, involved several steps

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.331486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.117552Z digest=sha256:40150fcd1c946d3b72f1acfd90c3036cc39199ce6551af09f0bfba71cc7f66a7

Observation 16d604f0-777d-47f1-8ccf-fb2d0fb1ceb4 · outbound

This paper cites description of overall motion.

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation description of overall motion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:01:57.318405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T22:01:57.121114Z digest=sha256:8d9bf1370a91300e5fa159191ee52ea406cf4f942306c70329cc38afc9981b75

Pith citing papers

Observation 988a81d2-fcf8-43f8-b760-b76e2c8df26f · inbound

Seeing Voices: Generating A-Roll Video from Audio with Mirage cites this paper.

Seeing Voices: Generating A-Roll Video from Audio with Mirage Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:17.425943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:17.425943Z digest=sha256:bd79aa0586c4de83647d9e16f01629cb4629bb2deac89030f84a2ac4336d3c35

Observation 22e46e91-26c5-42ff-a16e-a8bcfdf52b99 · inbound

Evolution of Video Generative Foundations cites this paper.

Evolution of Video Generative Foundations Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:05:51.761473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:41:38.616611Z digest=sha256:4c0ae143d39c3f82ef2daf7a7920b2932a7970842d5be9a16eed50e581f5eb4c