Pith. sign in

Paper Citation Record · LEDGER

MotiF: Making Text Count in Image Animation with Motion Focal Loss

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.16153.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16153 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:49:06.719292Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f8feaa9-cf3b-43e9-a30c-c2facc3da2bd · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.451901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.451901Z digest=sha256:797745baf3a65cd11c9f9bdb868c524f2c087a41c7a3db42045d0966e050c81c

Observation da23cae8-7d2e-4b07-afec-6a12ae46f14f · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.458641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.458641Z digest=sha256:d6c192ac1ba0482d7014be715c261c562ab0761243fd597de9cc81b81b1ec9e1

Observation 74fbbf72-d4d4-4804-86e7-bebc7489e8c6 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.463622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.463622Z digest=sha256:ffa7f160f5a3715fc17c51b103cd743fe937a9027bbc738345a0a8601c45c123

Observation b9e03d42-cbe0-4baf-a636-f1f8e651acf7 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.469399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.469399Z digest=sha256:f823e0de9f283b51cd3c394aa59baa2343f78698fded758adc68e18efe16dd73

Observation 4bf75c4d-7f4b-4156-a00c-cea98946bcc6 · outbound

This paper cites Video generation models as world simulators.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Video generation models as world simulators

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.474602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.474602Z digest=sha256:ee165ad15b28d7b846d84edef8b4ce1a56a8f33c747f203094991eca34285006

Observation bf199bd8-4545-4d3c-8709-440476e7bc59 · outbound

This paper cites Animat- ing general image with large visual motion model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Animat- ing general image with large visual motion model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.539067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.479805Z digest=sha256:ce47dab436e85604d2888e4197980e30fb54e76efe4446092a9036c8c841089e

Observation 32733eb9-4450-4ecb-b3f0-6df34be1be7f · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.484693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.484693Z digest=sha256:d7dbcff751ac9bf951fd8f5ffa04f620c5fe4e4761e31a6f401bc544641df131

Observation d54ab77c-b092-421e-a38a-7ca45abfbd3e · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Videocrafter2: Overcoming data limitations for high-quality video diffusion models, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.519339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.492435Z digest=sha256:469d3995b5a3f9942fe15cef591fc2028967bbbde3268584448218b818f0e672

Observation 7bce2c6f-c0a6-4870-b07e-767fb72f33a2 · outbound

This paper cites Seine: Short-to-long video diffu- sion model for generative transition and prediction.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Seine: Short-to-long video diffu- sion model for generative transition and prediction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.493575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.497289Z digest=sha256:77f628bb880a6f7a773d7262e523439699b4f2b0793889764f6de33083ad382b

Observation 17e117f3-301d-4a44-baff-cb78e7f84eef · outbound

This paper cites Livephoto: Real image animation with text-guided motion control.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Livephoto: Real image animation with text-guided motion control

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.473555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.503127Z digest=sha256:22a79c1123559f7c6baa557bf3d2f91821f34633b7be3e9ecdece3f54c71dbed

Observation f4ed73d5-72b2-471f-8811-223aa048cb00 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.508047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.508047Z digest=sha256:c1c2553424d75e03d6eecf407a0b371cbaec61fd075fd41ae25e9248904fe7d1

Observation 49397f0e-404f-4899-81a9-5da62a0eb0a4 · outbound

This paper cites Animateanything: Fine- grained open domain image animation with motion guid- ance.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Animateanything: Fine- grained open domain image animation with motion guid- ance

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.455046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.513281Z digest=sha256:f34c545119726f4175365d10b859d13ca92900209e164dd4b3576e7f18c0279f

Observation 0627cba8-ed4e-43c1-bd08-dc47d10d7ccb · outbound

This paper cites AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI.

MotiF: Making Text Count in Image Animation with Motion Focal Loss AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.518012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.518012Z digest=sha256:8dc5425f6716169942c726f7098de36b53b63cd9419b74c8f6d468a230098d58

Observation 88a32a85-c904-41fc-94e1-e40979e04783 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Preserve your own correlation: A noise prior for video diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.523213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.523213Z digest=sha256:ebdf697e89fc3bc6ddd8662684ab1f67d5a815cfe58f6b98897ab78b877bca3b

Observation 586eaac0-db1c-4445-b16a-69e6df17719f · outbound

This paper cites Emu video: Factoriz- ing text-to-video generation by explicit image conditioning.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Emu video: Factoriz- ing text-to-video generation by explicit image conditioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.428994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.528000Z digest=sha256:984985e6db682191194e0a943859882ad8bda4d95b412b733623ae88b18356d8

Observation 7955499d-8107-4545-9870-6a393211c43b · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss I2v-adapter: A general image-to-video adapter for diffusion models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.532316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.532316Z digest=sha256:3ea59badcdae0d05ebf6d503d81a6f9f33e0637f358f320783cf524e3e4df5f4

Observation aa24b8f4-5bac-4bf0-8c80-dbabc9fbd52c · outbound

This paper cites Denoising dif- fusion probabilistic models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Denoising dif- fusion probabilistic models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.536820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.536820Z digest=sha256:4e91500daa21c850c52085abbb93878ab265a53d4f6eb580052f679f5f324fed

Observation 560fb056-dcc2-4556-998d-8a1405e52b7e · outbound

This paper cites Video dif- fusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Video dif- fusion models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.541111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.541111Z digest=sha256:d04362d00c66a7dd7a14f2e2983df78abf73c47b1139b6f1986ed1b8a61a78e2

Observation d04e818a-69d6-49fa-9824-370b937f0d58 · outbound

This paper cites Make it move: Controllable image-to-video generation with text descrip- tions.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Make it move: Controllable image-to-video generation with text descrip- tions

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.375401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.545065Z digest=sha256:8567c8ddfdf3941486075b46cc86635c6d2b6c037e830ad9ce8567948bc9d926

Observation 12675c13-be6a-4cb3-bc47-a716c29613bf · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VBench: Com- prehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.358045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.549104Z digest=sha256:43ea53bd472ae7635f13848f6a1f2c0966b62b0683897f9a39e9cfba8f540f87

Observation 198c63aa-93f9-469f-a78b-ad352445e59b · outbound

This paper cites Vbench: Comprehensive bench- mark suite for video generative models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Vbench: Comprehensive bench- mark suite for video generative models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.341698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.553667Z digest=sha256:1276a6c4fcbdb71d17f933b4fc8ea132687112851ce8f51fe8d779d93a98f259

Observation c1e73ecd-7829-48da-bb5b-4573dde9b8d0 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.558534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.558534Z digest=sha256:13b6f8bd0e91adeadd1440dd39b2dc276b5d829c4dc6ebb270b4ab79c55a4b00

Observation 343e7d14-1382-4eb7-a5b8-5cf077a2a8b9 · outbound

This paper cites Physgen: Rigid-body physics-grounded image- to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Physgen: Rigid-body physics-grounded image- to-video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.324414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.563595Z digest=sha256:8261ad36cdacf927caf2f3c4fbca61a61487102590b797965a362944d20b55ef

Observation ffc46977-a43e-49ae-a806-a95416d3611a · outbound

This paper cites Evalcrafter: Benchmarking and eval- uating large video generation models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Evalcrafter: Benchmarking and eval- uating large video generation models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.567962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.567962Z digest=sha256:039ee30a4a751adc18f9c2a863cd5bbed9df36524ba9367f89e22dff3f4be0a3

Observation a7fca664-894b-4a25-8566-d6cff4a41ead · outbound

This paper cites Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.573171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.573171Z digest=sha256:e0d4375eb05030851d017da1b30d86a35307f4f2ff147434bc422544f7da741a

Observation 75213b1f-7a3a-41b8-b899-4ea7df769d5f · outbound

This paper cites Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.578398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.578398Z digest=sha256:3eedc6f161c68e7a709d104ff53b3abbc111cea4f0c121f6e037629995c69fee

Observation 5c4a0bb4-37e9-4cd9-b919-a87fdd65bcee · outbound

This paper cites Sync-draw: Automatic video generation using deep recurrent attentive architectures.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Sync-draw: Automatic video generation using deep recurrent attentive architectures

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.295710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.583317Z digest=sha256:1c18891166f73fbad2b7efc77b16d4c50e0ab21c8de185e6f3afbf37270b4732

Observation ce76bf80-9736-453e-bc6a-cf06adfcfbb1 · outbound

This paper cites Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Ti2v-zero: Zero-shot image condition- ing for text-to-video diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.277813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.588321Z digest=sha256:f7f26b805724287486c2be05177bd2f12f871fb4da1f343792f527e138d85d69

Observation 1ee9b71f-6541-4f7d-b6e0-ea914530cfaf · outbound

This paper cites Scalable diffusion models with transformers.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Scalable diffusion models with transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.593227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.593227Z digest=sha256:7f0d6471376208a97adf1a3f84de66136e005becfa8a4392da532f18bd791199

Observation 3b34f342-e379-4337-8539-e45aa07638f1 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Movie Gen: A Cast of Media Foundation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.598188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.598188Z digest=sha256:4adebd61d91bb24059d12cbb69f427f6ed42092ceea07ddfb6d8adbd09671326

Observation 0c3b48a3-4cec-4060-935e-db410c87f3c0 · outbound

This paper cites Hier- archical spatio-temporal decoupling for text-to-video gener- ation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Hier- archical spatio-temporal decoupling for text-to-video gener- ation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.603898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.603898Z digest=sha256:0f69af8c389f4485056a01bc06d8b22d9a50348cc3979114cf4b0590ad12e769

Observation 456fba28-d5e7-40d7-a7c2-35b61bdecc12 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MotiF: Making Text Count in Image Animation with Motion Focal Loss SAM 2: Segment Anything in Images and Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.609625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.609625Z digest=sha256:f89c5f32dbe3534772a97ed3e7cee7220043a25cd4a29d15c75f5a5c4e410262

Observation 1851ffcd-1ad3-4806-8e0e-b75bdf7dc961 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.615135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.615135Z digest=sha256:60f8977720ea09b809bbb50c8f31fc40cb18536905cb344d0bc82630a6319cd0

Observation c855f6b7-85a3-4d4b-944e-4428cd9f4637 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.620722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.620722Z digest=sha256:14c85686905e55829c6faea64db3e9af09e481e7071571a77f7e660cdc7bc17d

Observation 95ab1552-a16b-4c80-b701-d3549cbb9a92 · outbound

This paper cites Focal loss for dense ob- ject detection.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Focal loss for dense ob- ject detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.625570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.625570Z digest=sha256:75e29269d0a0dadbb2c33ef5d2f682370ae6514ba42abef9a53f8e340b7a20fd

Observation fb6607c7-825f-4d07-9c73-fb26afa6bdc2 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Progressive Distillation for Fast Sampling of Diffusion Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.630409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.630409Z digest=sha256:c6ce106c455907f650e119d58e3703b359960731b83d1cde9b8828afb1b12866

Observation d27b60d5-83a2-431d-9a44-6b2ee793c7d2 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.635388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.635388Z digest=sha256:c5cef06fca96e81ff0b542b9048ba9b1a09d159105f194d193e19b3cdb77831f

Observation da056e74-be47-4220-b25d-961c3f577954 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Deep unsupervised learning using nonequilibrium thermodynamics

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.640425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.640425Z digest=sha256:dd47137e7892cb4b1c0c2e5e03c8ce98d3f4862a747b2594e2c3c13982f67602

Observation c621ec9d-17c4-485e-9852-666146dbef3d · outbound

This paper cites Denois- ing diffusion implicit models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Denois- ing diffusion implicit models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.645239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.645239Z digest=sha256:aabc0b9ebe6df592fc1df340c2dd333bcff122270c0e534571caa4cee3fd3fa3

Observation 5466c700-6529-41f7-b510-725dfd94d289 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Score-Based Generative Modeling through Stochastic Differential Equations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.649998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.649998Z digest=sha256:6c4fde89091e566b1a18938a929ad518b086b4ac62473c39af6befa7941fd74c

Observation 4618c896-8899-46b6-a223-673af3e2432e · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MotiF: Making Text Count in Image Animation with Motion Focal Loss UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.654412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.654412Z digest=sha256:35407ad1f9864edcfc75d45f38fc7f9da3abb7082d8bce2768fa04ce41b01d62

Observation 9b8b7f74-b3ab-44d4-bd90-2016e500b1f2 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Raft: Recurrent all-pairs field transforms for optical flow

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.659181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.659181Z digest=sha256:f0b9eba6399bf68949a606985d3d21cb18924308c84dc49887f7f5181a87927b

Observation 78a763d8-6539-4f4b-a808-3efc8e5682ad · outbound

This paper cites Microcinema: A divide-and- conquer approach for text-to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Microcinema: A divide-and- conquer approach for text-to-video generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.187114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.663915Z digest=sha256:9530e40bb0884ce7901a6384956de30c385d84491a0c4b708655cedd79539132

Observation 44c9d339-6e15-4eaf-b087-3df805681924 · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Dreamvideo: Composing your dream videos with customized subject and motion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.668128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.668128Z digest=sha256:8dba9d8c9cb4300a4ff89dfd63ed9a6a2904ee1eb7158f8372582c737decbe61

Observation 3f8302a7-dfa1-4c62-8f72-61e94fd9df06 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.672568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.672568Z digest=sha256:5444428ac2daf65e2e9d2b8e555f9e41418a23b1bd00c0bb0c39ea9379704f6f

Observation 0ceb8851-790f-4283-8b4a-c5c16195bb4b · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.149060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.677016Z digest=sha256:83a2be5bc99cf01228074e53fee70797f98e8baf787bbbcfa1d1b36a564fa3e5

Observation 07a1458d-3778-4676-8bc7-48a26c7abdd2 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Msr-vtt: A large video description dataset for bridging video and language

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.131692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.681700Z digest=sha256:671579ef499b5df45ac1a0088c6719078e383d56736384670b413d438363d0af

Observation 5cb96f43-0e99-4955-a58a-98b8842049d8 · outbound

This paper cites Motion-Conditioned Image Animation for Video Editing.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Motion-Conditioned Image Animation for Video Editing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.686047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.686047Z digest=sha256:d130351d34fab4f0d9de535bd5bc72ae4a12e96b6e02f83f5464846d5fb4f350

Observation f3641f0b-fa18-4d10-9283-0702a1bc8f2c · outbound

This paper cites Zero-shot controllable image-to-video animation via motion decomposition.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Zero-shot controllable image-to-video animation via motion decomposition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.113704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.691065Z digest=sha256:199e53ded574971da14aac332de5865efa8efb67d1270284934c612d235b8543

Observation f28483e1-c41b-4d32-a134-fc6d89f71144 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.696189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.696189Z digest=sha256:2c0e97c6fdbc85aad18a6b58aa99d0ddce1266130e78d8ca17db9fa2afddae14

Observation 168f8b80-745a-4b15-995b-140bdf2a3a62 · outbound

This paper cites Pia: Your personalized image animator via plug-and-play modules in text-to-image models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Pia: Your personalized image animator via plug-and-play modules in text-to-image models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:49:07.095893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T10:49:06.701883Z digest=sha256:4298ec3588df88ffafcde5bd66b881f5dfa8a495818884c4127da1009a57c509

Observation 0f481d93-ab49-4145-9808-ea9fe352c893 · outbound

This paper cites Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.709009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.709009Z digest=sha256:6705b797ed0e0acd43399651559017651cda991ad695863b3357034a5eef4a16

Observation e25f6935-7bbd-4cc5-abc3-6af1b4c76252 · outbound

This paper cites Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model.

MotiF: Making Text Count in Image Animation with Motion Focal Loss Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:06.714131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.714131Z digest=sha256:5dc511f6df31922861bd6e990b37344fc2fd4d960bcbeb3092866e14023555e1

Observation 2672eec0-28a5-49b0-943f-c08c0e4192e4 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

MotiF: Making Text Count in Image Animation with Motion Focal Loss MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-11T10:49:06.719292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:06.719292Z digest=sha256:9ab2117430e30ac91ef3fa315aee86048927a9400662351e2445b3b915a0c88a

Pith citing papers

No inbound Pith citation observations are available.