Pith. sign in

Paper Citation Record · LEDGER

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

As of 15 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.13753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13753 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0aa1032-d329-4d63-8247-9bb45a30bd19 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.327857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:03.661367Z digest=sha256:0da1ef0aaaee886b71d37a7544d2099fc86e76627aa9e96b513fe9a937239da0

Observation fea59f5e-5851-4875-abd1-90ec627f9a3e · outbound

This paper cites Huang, and Niloy J.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Huang, and Niloy J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.308915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:03.740191Z digest=sha256:448baa0e77b345d409fbcdaf35842163b365a8eb2033494f969fa2462b502cec

Observation 3dfc697b-0294-40d4-abc9-6684c7e55d45 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.293233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:03.826011Z digest=sha256:402702f060141fab8df0b6638a7035578e3d72dce4fa7a3a3eab710db21f565a

Observation 0e1bcdd9-2d98-40f4-874d-dbd2d406b9e6 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.276753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:03.932309Z digest=sha256:99e5d57a3024227f444688d228fe57dc47576cece72f244cad3d435e4f935138

Observation a87cc984-5599-44c9-9934-d66f5bb13213 · outbound

This paper cites Diffsynth: La- tent in-iteration deflickering for realistic video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffsynth: La- tent in-iteration deflickering for realistic video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.261382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:04.026929Z digest=sha256:8c33a71d9ed4e670fb1cb86449b1d9d16001a15253b3d66a2bc0fba41b09fdf9

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:690526abb3800eb8d1cd8f6625a3fadf8208c574369ab94e6bc70c379896036f

Observation 9072bd9c-9ee6-4a21-9ce7-a2fabb2afd5d · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.195327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.195327Z digest=sha256:381adbc43430eb1203537e06e7d3995f8a14cb0a866d3c75ef4281ad61682acd

Observation 198ec7a1-a731-4f2f-b311-514e79b0cb7a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.284760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.284760Z digest=sha256:a010f53422e677d096225795a213db833f7c9b747c833399e347d6a895c7e31d

Observation e8be94f5-25dc-4c24-b133-79ac6cb494f2 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.387504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.387504Z digest=sha256:164b6ae2aaabe8db686d0a8f5aecbe6b154cae92bfa0d9b344a3a8ca93ebdd34

Observation 4c6848b5-8a91-4dd0-b625-b2ed812a968b · outbound

This paper cites Video dif- fusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video dif- fusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.473396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.473396Z digest=sha256:6edc13ae9627ff952bdae18a9403f0441020426ca9c792794ee8c96995216214

Observation 392234db-f6f1-463e-9773-724403ced627 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.571398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.571398Z digest=sha256:1661315efb0129e510d817f34b852e46ed8868535f88d0ac1b058493b0dfc762

Observation 2dc5c1f7-20e0-4835-b18d-5dbda1226131 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vbench: Comprehensive benchmark suite for video generative models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.073040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:04.658315Z digest=sha256:320d244031d2037f465128c88b7377c9fc9138f72c820df712e0261fcbb28066

Observation 9bd42f93-9578-4d15-8139-3352f5e9e531 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.055169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:04.779140Z digest=sha256:8dddac63ce463d364b9e601ccf3a6b1fd5028d0b638e1cbc760aa098165fd5bc

Observation df7d391b-9067-43e4-85d2-c0e13ece3784 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.867561Z digest=sha256:d88c5be1070e27361d5033544182c4aa8983290a0f6a717b49424e20a742bcf6

Observation a9ef22ac-c314-4987-9511-ed90bea7f57e · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.956206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.956206Z digest=sha256:1893b38fc354a93a4d0fdba270ae339d25e2aa767217e4275eb2d3eb1bf286ef

Observation 29fcdf9c-5ae6-4433-913d-484fa50812f1 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.030716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.075194Z digest=sha256:a37b2d3da6c5c93e9b0c5987c5caeaf2c58b02fdfa0979101a7eb912adec5212

Observation 6ba76e2c-012c-4315-96b6-8d829a3a266e · outbound

This paper cites Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.001903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.139965Z digest=sha256:8f16e335830dee3372fc3264b84c86b3891ba77b1b3c9fd3301f0487c15ba346

Observation 36f84924-0470-431f-8eb2-8773afda9bb5 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Open-Sora Plan: Open-Source Large Video Generation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.217568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.217568Z digest=sha256:e81f2ba6abfce5feead0355de87a9498f2816f48224eeb5ba63a5195d7bb695f

Observation 8fa1c850-ab78-40ae-a79a-833be2520ff8 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.312282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.312282Z digest=sha256:65772c39b2483732c744e601cb8730ea88b1d8cae7c4ee2489099bc2e97376d5

Observation c554a101-d293-4e2f-bd43-d2de26e0588c · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Towards understanding cross and self-attention in stable diffusion for text-guided image editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.982327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.379869Z digest=sha256:62b69daaf8574c96df3e9cbcdc0cef863bba87d2eed4b176a3fdab89f537f9ee

Observation f9958775-c95f-4fb9-9431-f61f11ac0c70 · outbound

This paper cites At- tentive linguistic tracking in diffusion models for training- free text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis At- tentive linguistic tracking in diffusion models for training- free text-guided image editing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.956632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.466037Z digest=sha256:a99235d29e7daea65f83e2e30db1c3f80415edd632ca5817b0ff2fa33d26bfa6

Observation e0d2149e-da78-48e9-9e14-de0bfcdc9884 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video-p2p: Video editing with cross-attention control

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.939679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.555043Z digest=sha256:0eaac2a7ff322825337afce1a2b7e7d4ef4f19bae063cc6ce20ff501ada7a82b

Observation 918aafef-6849-436e-92b0-5db97238a645 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Evalcrafter: Benchmarking and evaluating large video generation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.625822Z digest=sha256:ad9edfefcd4855bfde23d86a5a3b8285db9de1bd930678fd97232b9a1d0d694b

Observation a4c8f94e-c403-4eb8-8480-bbe6e462190b · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.715742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.715742Z digest=sha256:f5db0553c0676317b6b64817c5cf59ccf39da024c8021ad96f6eb5dfd65dcbd1

Observation d72f10aa-8b15-455e-8fe0-e02e0c29693b · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Null-text inversion for editing real images using guided diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.897595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.801114Z digest=sha256:e7e3eea1df70b95614cfc874ee32b0c8594c124b74c15ded6ca9b672cc0c9b31

Observation 6d940d20-b81a-42e8-97de-901cd1f70086 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffusion Models for Adversarial Purification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.886229Z digest=sha256:f6b154a558067e981e60a7da4949f90392871ee7c60361dfe2c9cfca610941a0

Observation 6473f36f-99c7-45fe-bfb0-107d9b29ee7f · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:10.647955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:05.967881Z digest=sha256:bffa2fd292e7b6d627f7780463c23dce4185b9b7e5aa05c5e871782452a7477d

Observation 15b108a6-e171-4e04-8766-8e853205d8b4 · outbound

This paper cites Pika 1.0.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Pika 1.0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.396746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.036782Z digest=sha256:29832148d375765ecf2597e656c82003822da8420f933c3e378d4f18a615b378

Observation f3e0ee66-6b11-45e7-99fe-3ba01e1f947e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.099876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.099876Z digest=sha256:abac960be3e3e336b447267e6b0ff5d9e55739741e974706941dd6b411fcd7f9

Observation 375ff990-6126-4afb-a6eb-1ae9bb0c884b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.276468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.189085Z digest=sha256:b515baeaa5943ed5d5a95131e144c75cd28adaf75ef852ae451a447edc183193

Observation 76492959-6da0-4ef4-9a55-911f3565b125 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.087093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.293597Z digest=sha256:f4d2549fad745182a2a0e39d92bb61b708c15670ce51cb880f55ea2f89f9a933

Observation 907c4bd1-90d0-4049-b627-4946d099121e · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.833297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.383786Z digest=sha256:2d511d44101f85999a9e64fb8fb0d9632f59d2b931f68deefb0929aea1bd9066

Observation a3af56a6-1e74-46b2-95c1-85849e106bab · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.717542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.484434Z digest=sha256:6a8e80cdac6392a149cdd7514f0bbd17617f0026ce742e9344a48f3a2668378c

Observation 320ef7f0-0c21-4df9-940f-9781f83cfd51 · outbound

This paper cites Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.612955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.570787Z digest=sha256:bd8de56185cd6a065c54d62a949ffb1c9b89f36dec05640eac1d6454d645667a

Observation aab08ec3-8f12-4e8a-9e98-8016e78cb9fa · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Edit-a-video: Single video editing with object-aware consistency

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.422686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.666069Z digest=sha256:da2c36f018e671bd3ccec15358ce5832de9995a2d8d2e397301e13b5a199cd36

Observation b4988830-1ec9-464e-b471-d6a70db25a75 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.746047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.746047Z digest=sha256:325d2cd30738f76487acd23be0546fa8f3b76fa4ee0741da53a4422753b70be0

Observation 328a147e-05c2-41d5-81eb-ab35e98dfc4f · outbound

This paper cites Denoising Diffusion Implicit Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Denoising Diffusion Implicit Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.829645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.829645Z digest=sha256:572df6662251ba3c3f3dcc45875947eb14662ccf662069598b987360cb28461a

Observation 8657e463-d5c8-4a85-913c-bcf6d3112a97 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Plug-and-play diffusion features for text-driven image-to-image translation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.249295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:06.915722Z digest=sha256:6feccdda1c4503ee4bd5be945e3518f19c4219c24843f605ff7018e2b3ff7d2e

Observation 845dbd66-6d58-4869-8af5-3aa2b0ed9879 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ModelScope Text-to-Video Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.980555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.980555Z digest=sha256:36170c31383263edf3f0c9ff8a373af8abbcfebb5f606e5f4461d41ebd6c91d7

Observation f6e5d777-157d-4421-9480-31be1a41639a · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.087178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.087178Z digest=sha256:04f9789db6c67f9d796b522abbde1862a614be12a4afee93112d621ce20dcbff

Observation a9a1aa62-1b7b-41d2-9abd-0de5e46b7720 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.038369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:07.174284Z digest=sha256:c1e3b52aeb4e9b97a70dca1e5ea075c1609d85d712d857c0b350c98479d10d5e

Observation 7cf0cde8-aba3-4dc7-b04d-146c968e51e1 · outbound

This paper cites Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.865062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:07.254277Z digest=sha256:73c97f6c11642630cc051f04b1fe400b8722c5a686f42d9d7f98c850aa7578b2

Observation ffa0efa5-6e01-4291-81c3-745e8b7e73d2 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Rerender a video: Zero-shot text-guided video-to-video translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.669898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:07.339187Z digest=sha256:169eae65b56d4a178a4762b04c203cbed5b5b832dedac3ae0ff558241c823c85

Observation a13928cf-a1bb-4530-9664-af5f101d3778 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fresco: Spatial-temporal correspondence for zero-shot video translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.509404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:07.427186Z digest=sha256:0d89904916829f6b1b94f8e3837966b8da938aa16d74338d7e30c6fbce9d2cac

Observation 60f5666e-ec18-497c-99ac-9f1789767b36 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.516251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.516251Z digest=sha256:5b164a759eedfbd57a5637063d3cda28fa2289af96722a4fabc83b533370130c

Observation 97fe92c6-dd19-48d1-8024-14ca4544c834 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.602629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.602629Z digest=sha256:f571dfbce9a101ea5843d77c9c177df07422410948f30a859fac0e7842a39f43

Observation 6f4480d7-b3f2-4914-bd8d-0eb6e57c9371 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.699049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.699049Z digest=sha256:a9108f673d558704821093de7dc6413be16d97aee11c5c6d9cf044a19f0cd444

Observation ec343ff1-a59f-4dbf-ac32-4a867b6a9583 · outbound

This paper cites VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:24:08.016744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:24:07.779000Z digest=sha256:a7cabcce863996952ca00a045ec7d946018ec7c98ecae5e8b01ca369aecc8434

Pith citing papers

No inbound Pith citation observations are available.