Pith. sign in

Paper Citation Record · LEDGER

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

As of 17 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2501.16714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16714 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:17:24.903509Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:25:41.219753Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T07:27:09.000021Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad7f1e2a-4252-427b-a6df-8b5776f895bd · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.716992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.716992Z digest=sha256:e46780c597901b5241e351ec84c06a228de2762dc562d3cd9cc2060a80c9a315

Observation 61ceccf2-ca3e-4887-a4b5-29e859291702 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 2

Resolution
verified exact
doi, observed 2026-08-10T11:17:24.935013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.721436Z digest=sha256:5508c2e5a68af32a09b54b6a2b2245e2a47fc7131d8489ffd17c85cb73197630

Observation a6e6128d-7135-4657-b0a2-29aab04e3505 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.725361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.725361Z digest=sha256:a545182ada83a7ba149963713d17fd82a2c1b39da76902bfde3f33a7eb2ab4bf

Observation 8d118574-d596-42b7-b080-f52397ffaad1 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.410757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.729581Z digest=sha256:47ff1cd8215dd65568caaf7a4ec398c09832bce37501db1c8f0acac53fbf44a6

Observation a0481817-f1df-4f05-9cfd-c858b7d0c771 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.733141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.733141Z digest=sha256:d408be3ce5ba146750c79809857e2ead74f5bf6b1f919ae2dc2f91188eb731c5

Observation da6fc8be-d217-489d-9261-d89027a9fb99 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.393061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.736786Z digest=sha256:b4f421c09459a9e2f9c2145d471cb08cfe6fa97e24702a3d28646fa80a0f96d9

Observation ad45eb0c-e1c2-46d8-9716-2eb1cc99efbc · outbound

This paper cites Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.740435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.740435Z digest=sha256:2daff1b8e8e8353c840e65559e6a0d5194c68f4ae33a566fc69067ea7a2df2d4

Observation b32d2c54-5cf4-4e9b-b5d4-a229e7c43a1a · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.380908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.744358Z digest=sha256:fab214fc666145f9402c037e9616bf429a08baea2680e128e6dc5efc0f9f12bd

Observation 1304b929-9e51-4aaa-950c-0cb567d6c633 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.369699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.747808Z digest=sha256:1c7b43b76a087e3ebc0412e2848f4134d7f062f5a74a645e2359748410813e94

Observation 08a1bc57-c39a-47f3-943f-1c15b1c1d9e3 · outbound

This paper cites Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.751477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.751477Z digest=sha256:68c6288ee9e08e7ecfebba18356049385d16c4770b745879df40a76b9e11ea4d

Observation 340180da-2218-433c-9309-841462e264b2 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.755854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.755854Z digest=sha256:071b7ed0ae59c31b0679ff23a030294bcff97307b9b0f75183127c498072eac3

Observation 445a7ea7-14ee-4961-b479-f35071e3cd39 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.759958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.759958Z digest=sha256:644b8c16eeb4911de865ff6b85eecd33db0f9f8271de14489ccdba36c81b74b4

Observation bc0adb3a-c2a4-4c65-a64d-b534cceaa7e3 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.763980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.763980Z digest=sha256:eb7371b241f8034a00078cc7fb652dc3107b940f9d58c4d19945f4e87f298e71

Observation 97a54076-6c1a-4e8e-92c3-dfd00cccd9b3 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.768234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.768234Z digest=sha256:fd91b57d27b4520ff7ba1db49a88925c9143697531ca4523f44947f92ff6665e

Observation 0048a63c-e256-4ea6-98ce-3d561b9747d6 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.771671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.771671Z digest=sha256:d7c7856a3e601a20d37344a127c34a00d8ff7a47ccbb0420b2d96adc8dc9cf7d

Observation f69ca729-401c-4ee7-866a-fc22019495d0 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.353317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.775620Z digest=sha256:8d334cac16a951e97eba6cceb09ddd1171c8ce087e3b19e7c433d6889cf62655

Observation 87adc961-1d8f-4e2c-9c29-d6893d273a19 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.782768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.782768Z digest=sha256:5fd681ee6d51d49bc471797d48caf03cc2b2a9541ed30c2fabd9f21b031c568a

Observation d0f7e956-d73b-46a2-8707-6d72d2e3995b · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.326812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.786865Z digest=sha256:ac97e28ff8dfee6c5ce38bd343d6e8d6cc37693455e4e7291357dbb443c84695

Observation 9b90ddee-ad18-43ca-8a26-de9d0a18f955 · outbound

This paper cites Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.790436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.790436Z digest=sha256:10078fdb60a0e460900cee47b77556c3097560a89e2c31da4e5e3423b18cc0d6

Observation 93464f84-44f6-4507-90c6-60d99139d605 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.315972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.794465Z digest=sha256:c8e5525644c573cccd8e65474631b478c5eacab3771d5ff4991ad1bdd6ec5c6b

Observation 3b18f5b5-e233-40e2-978a-75b4ac4d4346 · outbound

This paper cites CCVS: Context-aware Controllable Video Synthesis.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models CCVS: Context-aware Controllable Video Synthesis

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:17:25.081647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.798016Z digest=sha256:6c65c70ebda85510a382b47be37abf1578f5f9653af4dd1e6f14a540ea0328bd

Observation aab212cf-d9a0-4e96-9d0a-4c82f379aa1f · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.801805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.801805Z digest=sha256:af5a47398101689e5292adfde259265e5bd77899b4bf3f2c85736941fd0b29da

Observation 943063ba-1ee3-4462-a2d4-c0df3570d0b2 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.805507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.805507Z digest=sha256:4637661b9ad6d1284f251a50405e0437e63f2a875fc4e9f4ebfc049a21b015b0

Observation 029edbbf-bb42-46a8-a2f0-8e4ff213697e · outbound

This paper cites Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.809060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.809060Z digest=sha256:c841668ba631aca8bf18d77d89417066c793620801bc8c4a499b62a6a3a218ee

Observation f4e875c1-b4a7-405d-8c7d-9390a7ad85f3 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.289508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.813104Z digest=sha256:5aad5e384415b66664f6a7e25c7538b6599ec70b7326ad6288bdec42a3973d04

Observation f405fccc-477f-4fe8-9d92-7776899d0312 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.816691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.816691Z digest=sha256:693408726e5874934269a1f5b33ab745f9be8e12feedb26b5ad0a1128d635119

Observation 41f599ec-a67f-4ea5-b07f-c202c896b9d1 · outbound

This paper cites MoStGAN-V: Video Generation with Temporal Motion Styles.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models MoStGAN-V: Video Generation with Temporal Motion Styles

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:17:25.055305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.820268Z digest=sha256:f232660e10ecc21962671ec35c310ad78f40322faf2a4636182c98cf93858e0b

Observation 38f3c9ed-7dd0-4c0d-b679-21a460b60043 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.269787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.824388Z digest=sha256:a215949d4ad63d22ffa4c5f15ad104f3129f689aba4dea940ac2cba70a85996e

Observation c12dba9d-7315-4e93-a35b-123df795f689 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.828362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.828362Z digest=sha256:146fa9fd47abba9cd8334253977e7eb0eac0e369189065e01068c9be45cdae87

Observation 8dc7b40b-655c-4f67-aa10-21d93b08478d · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.259129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.832598Z digest=sha256:b0b04b530ade4442027489a6d7e3fbb682d422a173a6051a1691dad42c028740

Observation 7d5894b2-323f-4405-b314-fc4b857b2fdf · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.836297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.836297Z digest=sha256:e9171b07ae9cd00c578af4abb9e977c0d431715df4fea79b8611b8715377a5c8

Observation ae3acb05-92de-4a58-b489-53a06c048092 · outbound

This paper cites A Good Image Generator Is What You Need for High-Resolution Video Synthesis.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models A Good Image Generator Is What You Need for High-Resolution Video Synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.839983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.839983Z digest=sha256:d6a03a03bd09fda359bf81796acefbef6839cd5acd4dc992684cff347e021e5a

Observation a91253d5-e832-458b-8a61-6f89d920ac4c · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.843887Z digest=sha256:d96a952040aa7b19fd4e0cfbbbb9a12c3d0653647be2024c04f9be79bc47fc79

Observation 5fb778f4-fb5c-41b4-b875-20c4f455237c · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models ModelScope Text-to-Video Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.847643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.847643Z digest=sha256:cc6132258770412f9e0059bf1267798d691bcb89862ec08c33f1e54f8db19baf

Observation 7a5a5e44-345b-4c7d-ba34-13234c58a52f · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.233812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.851441Z digest=sha256:fa7a6992c3cff00bada0e6e1cb8b54538116723b8b0bdbc348fefaa2857b2def

Observation 4b653dc5-c7c2-4fc4-b10e-e1a3f159239a · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.856008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.856008Z digest=sha256:7b1947c6b3caa0c0632f5b9435c70df8aa78e2a97756b15a1e0002ff11f7564f

Observation 0f555894-55a0-40fc-a59f-49b23db7f74f · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.860443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.860443Z digest=sha256:fe3c7b011afd62e478e063608f7e993f857c6c145949dfe6c59ca497ccf80d48

Observation 6469d19e-a933-471e-8e74-db46f226ec1a · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.864384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.864384Z digest=sha256:472bc74c645dfecdf9e1f367910b8a339ac9fd11beb3275194c28f90bdb62c11

Observation ccd28ec0-cff8-4385-b012-030e68eef906 · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models CVPR 2023 Text Guided Video Editing Competition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.868298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.868298Z digest=sha256:76b241a18a52f02619f948564c05acb2a6f0fc79d9cb6905956048e9961777f5

Observation 59512b05-60ee-49a2-824d-0240f012b100 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.207227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.872630Z digest=sha256:b76f5ff3a86498b4adf9543660964606f6ed01cc08e2fb0dde7812ebf71bba5c

Observation 3007c7c4-68ef-49c6-b1af-1711ce76c993 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.195218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.876562Z digest=sha256:08ae75da54656f8a14cdb290feaf5fc293f7ab96c5b1d6817548b6be7f116a1a

Observation 32063917-0a22-4d9a-9290-6bef0109ab90 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.881055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.881055Z digest=sha256:8a984c38189d4322450d9ecd0b7b23376b83a668662b7bffcd3a0d9d9ede8855

Observation 17d21a95-98f7-46bb-b3a6-06823772285a · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.885047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.885047Z digest=sha256:8037fd366bdbab56ca8fad7170d952be8fdc3686581430dc15ab697adaaee1c9

Observation 501f2b6f-5abd-4022-a7bc-d9cacc3cced1 · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.888752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.888752Z digest=sha256:c0b0cfbd95f8cfcf284c3dc6fd6df22dfb3e286c7b18ebf857e347d974c39069

Observation cecb1ca9-6dcd-42c1-82b5-466deb25c6c1 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.892542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.892542Z digest=sha256:3a26a0dd452998f05e5a10f9dd369e4d7b6d3791b1af06b480a4a7ff839dd4cd

Observation 162719ec-b930-4feb-9a5c-4f9162495ce4 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.896071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.896071Z digest=sha256:f8c46853128fb7f713fcdd72c85e4fe68cfa9dfe35b28d2b25e3907aefd4b9d9

Observation 8cf55463-ba75-4397-8a5d-1ad2035962a9 · outbound

This paper cites MotionDirector: Motion Customization of Text-to-Video Diffusion Models.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models MotionDirector: Motion Customization of Text-to-Video Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.899736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.899736Z digest=sha256:a22de710d6dab10b0aff352c55bc0eb518e10abbad6310c54bca31e8183b0967

Observation 878c0bc6-71f8-4a08-96bc-4a05028269bc · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T11:17:24.903509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:17:24.903509Z digest=sha256:4c153371306201d2a584cd07b0f0acd010404e30ff7e8ca6b57714f405270d81

Observation 01051b4d-0ff1-48c4-a1f1-0e70cfbd5087 · outbound

This paper cites an unresolved cited work.

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:17:25.343350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T11:17:24.779416Z digest=sha256:40face3f25c45f4ba3bc4c1c091ea411d8aea4dcd505a935ea5cc5b754ce4134

Pith citing papers

Observation 5f29e41d-f228-456c-88b1-cfec2d8fa30f · inbound

GenHSI: Controllable Generation of Human-Scene Interaction Videos cites this paper.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.001999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:05485468d535fbaabd5c41f9e45db2babedcd56a1d03e2e9706c6f6de8f764b9