Pith. sign in

Paper Citation Record · LEDGER

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis

As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.13753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13753 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:24:07.779000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0aa1032-d329-4d63-8247-9bb45a30bd19 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Emerg- ing properties in self-supervised vision transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.327857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:03.661367Z digest=sha256:16cabf7d5dfb6f6288eb3605cff4c61188727c8e24efd8ec3e09ec291cc4e3ed

Observation fea59f5e-5851-4875-abd1-90ec627f9a3e · outbound

This paper cites Huang, and Niloy J.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Huang, and Niloy J

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.308915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:03.740191Z digest=sha256:c701d0f25132bf16f5d75704c0c14511316ff37d350826f0864e41a1607e2d21

Observation 3dfc697b-0294-40d4-abc9-6684c7e55d45 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Videocrafter2: Overcoming data limitations for high-quality video diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.293233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:03.826011Z digest=sha256:589d31bf50594de0a994467247392723216c0bf78a4ea319db37e5b78a467002

Observation 0e1bcdd9-2d98-40f4-874d-dbd2d406b9e6 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.276753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:03.932309Z digest=sha256:3fb537f28bbe0b8a8ffc32950bc24868a6fe21bac931b8d04f0f572f264db83a

Observation a87cc984-5599-44c9-9934-d66f5bb13213 · outbound

This paper cites Diffsynth: La- tent in-iteration deflickering for realistic video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffsynth: La- tent in-iteration deflickering for realistic video synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.261382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:04.026929Z digest=sha256:17abd88fea7cac81db24a500f78cd0590440331cbdebb9ea059694fe0ac1927e

Observation 5741b480-aabe-4f75-800f-eb9fc30f5b5d · outbound

This paper cites Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.099250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.099250Z digest=sha256:3e00a6e5993541a082bf9f9fe6f5dc32b8c7866ffb47361aac37dc1dad6a1ed0

Observation 9072bd9c-9ee6-4a21-9ce7-a2fabb2afd5d · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.195327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.195327Z digest=sha256:bfbd91618d403dea3eb1150803fe9386c15c8ff2916f4778f3f19de3ce60f666

Observation 198ec7a1-a731-4f2f-b311-514e79b0cb7a · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.284760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.284760Z digest=sha256:bd6c2a44b6a3fd5c564a37c7a0f8a928f8c1224a55ffbf942cc176871f1d9f4b

Observation e8be94f5-25dc-4c24-b133-79ac6cb494f2 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.387504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.387504Z digest=sha256:e8a6901c87bfde4e4a47955f09c6a5e4be202f7de2331ec4ddefd799cb0ca7b9

Observation 4c6848b5-8a91-4dd0-b625-b2ed812a968b · outbound

This paper cites Video dif- fusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video dif- fusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.473396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.473396Z digest=sha256:ca9cd6302e4e3d48d376e6d93c325074d463f1b0ddbf55b3fefbce80a419e305

Observation 392234db-f6f1-463e-9773-724403ced627 · outbound

This paper cites VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.571398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.571398Z digest=sha256:ca2fa707c0ee89eeaaccc495bbb0c0b9e5462b246676dc279bc958e977b97576

Observation 2dc5c1f7-20e0-4835-b18d-5dbda1226131 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Vbench: Comprehensive benchmark suite for video generative models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.073040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:04.658315Z digest=sha256:26b8aac60b34fe25ad1be9337faeeb6dfb63d16c864c88b3a9b97df58173ab5d

Observation 9bd42f93-9578-4d15-8139-3352f5e9e531 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.055169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:04.779140Z digest=sha256:a21a92f6d8798b4534a4e08313adcc8f7948c0075cef96ad7b70ff0aaa184af7

Observation df7d391b-9067-43e4-85d2-c0e13ece3784 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.867561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.867561Z digest=sha256:c4b2d34d9946e46cd7b92639e8ea47ae0097b72f326646e38a34041ac2da8ea0

Observation a9ef22ac-c314-4987-9511-ed90bea7f57e · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:04.956206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:04.956206Z digest=sha256:087102ab93b9b948d4c6dd4cf9c71b71e2dac1e90c1ce928b4c3d78263e88622

Observation 29fcdf9c-5ae6-4433-913d-484fa50812f1 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.030716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.075194Z digest=sha256:a33c1ebf6621ed43f6229ed1eaa0052d59358b65f4f64b5e3f2f9f3c549ef4f5

Observation 6ba76e2c-012c-4315-96b6-8d829a3a266e · outbound

This paper cites Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Flowvid: Taming imperfect op- tical flows for consistent video-to-video synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:11.001903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.139965Z digest=sha256:4360dfb7acfc3acc26770f08c19c7475673f07e21a33bb293a204ca3d2d3ebe6

Observation 36f84924-0470-431f-8eb2-8773afda9bb5 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Open-Sora Plan: Open-Source Large Video Generation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.217568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.217568Z digest=sha256:753b9559cbad9f28f210429db9e83d62a41e86b9134f2d9e290f9593436d6131

Observation 8fa1c850-ab78-40ae-a79a-833be2520ff8 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.312282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.312282Z digest=sha256:91f404b6f93c9f04cbff236b3d41f517edeb347640cef61d1717078f5a57f4fd

Observation c554a101-d293-4e2f-bd43-d2de26e0588c · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Towards understanding cross and self-attention in stable diffusion for text-guided image editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.982327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.379869Z digest=sha256:c70c5eea0cba9cf764b66e9821c746b839faf1b6bacb8379dd1221ba59b419b9

Observation f9958775-c95f-4fb9-9431-f61f11ac0c70 · outbound

This paper cites At- tentive linguistic tracking in diffusion models for training- free text-guided image editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis At- tentive linguistic tracking in diffusion models for training- free text-guided image editing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.956632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.466037Z digest=sha256:cb0404833194b518848595c01406317febcf58ab73af99c0799adab8c9d7087e

Observation e0d2149e-da78-48e9-9e14-de0bfcdc9884 · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Video-p2p: Video editing with cross-attention control

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.939679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.555043Z digest=sha256:5bb64b057230b7a6c37628404b790fcf520508d4bc84ce407a377819bb4e32ca

Observation 918aafef-6849-436e-92b0-5db97238a645 · outbound

This paper cites Evalcrafter: Benchmarking and evaluating large video generation models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Evalcrafter: Benchmarking and evaluating large video generation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.920453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.625822Z digest=sha256:c5955f27e5d46efe4357d663e484a72216177c1c5d5e05d6a03d8ff9c774cffa

Observation a4c8f94e-c403-4eb8-8480-bbe6e462190b · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.715742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.715742Z digest=sha256:80751425e0c9665568243ec19a065eae5c88c77361b77a59b072ea47a3947922

Observation d72f10aa-8b15-455e-8fe0-e02e0c29693b · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Null-text inversion for editing real images using guided diffusion models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.897595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.801114Z digest=sha256:bd319614772fdd482d10b7ff5184a4a95109b3715f733994fd4d9f25bf7d6d92

Observation 6d940d20-b81a-42e8-97de-901cd1f70086 · outbound

This paper cites Diffusion Models for Adversarial Purification.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Diffusion Models for Adversarial Purification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:05.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:05.886229Z digest=sha256:5fec26e1f1bf82a1b075349dbd52501517733ede82d3e1abded476c94f70d6d1

Observation 6473f36f-99c7-45fe-bfb0-107d9b29ee7f · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:10.647955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:05.967881Z digest=sha256:74ba3d9a258fe71986468d2bac1a3524391964a77054dc2cac2d12b211436429

Observation 15b108a6-e171-4e04-8766-8e853205d8b4 · outbound

This paper cites Pika 1.0.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Pika 1.0

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.396746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.036782Z digest=sha256:bd6382a8a3f8963196ab3644b29c61fc52c6109427f96c8d1963ccf6243354b0

Observation f3e0ee66-6b11-45e7-99fe-3ba01e1f947e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.099876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.099876Z digest=sha256:897ed5831b30055f27d89d6c1c1a8f57d5e1ac39a4df3369d201f9c689a4c2f8

Observation 375ff990-6126-4afb-a6eb-1ae9bb0c884b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.276468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.189085Z digest=sha256:ae88d29511f85db5d3a66bd512f34149de539f5124268b1e69a24a0e677cf603

Observation 76492959-6da0-4ef4-9a55-911f3565b125 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:10.087093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.293597Z digest=sha256:c75e12ba81861ec0b7bc0bde7b4aad8467180c8172c0d77d1b0a3f56a0cd4981

Observation 907c4bd1-90d0-4049-b627-4946d099121e · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.833297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.383786Z digest=sha256:3c048b9923e1e43b736aadac22cdbba6df8137641c3ebef59805dfeaa7786d66

Observation a3af56a6-1e74-46b2-95c1-85849e106bab · outbound

This paper cites an unresolved cited work.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:24:09.717542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.484434Z digest=sha256:38e0b4ac1452be5ba95c5fc24d4d2895f20603ea4fed31cc04c7981ff703010e

Observation 320ef7f0-0c21-4df9-940f-9781f83cfd51 · outbound

This paper cites Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Bivdiff: A training-free framework for general-purpose video synthesis via bridging image and video diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.612955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.570787Z digest=sha256:9f755b87b270f4df6a500b125f9bf0b55e319a2fdc8f60cd0a38946223dcf765

Observation aab08ec3-8f12-4e8a-9e98-8016e78cb9fa · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Edit-a-video: Single video editing with object-aware consistency

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.422686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.666069Z digest=sha256:74c3e5e1f2a25a37c841fc0db8d4c497f45a1b3ac61aaddd088732491cf61c97

Observation b4988830-1ec9-464e-b471-d6a70db25a75 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.746047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.746047Z digest=sha256:97c35ba947a34b94b569ed393b42ff888986a2ee0f38cb3d2ad90fbc367a6687

Observation 328a147e-05c2-41d5-81eb-ab35e98dfc4f · outbound

This paper cites Denoising Diffusion Implicit Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Denoising Diffusion Implicit Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.829645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.829645Z digest=sha256:a66082b665f72f0ce06cd3a03519fd0aab903f7a93e1dcf62ef586a71afeb9f5

Observation 8657e463-d5c8-4a85-913c-bcf6d3112a97 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Plug-and-play diffusion features for text-driven image-to-image translation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.249295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:06.915722Z digest=sha256:3091366aac0c8b75a5a5f69ba8ec0bdd47f787a242164098c91239d8749279e8

Observation 845dbd66-6d58-4869-8af5-3aa2b0ed9879 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ModelScope Text-to-Video Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:06.980555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:06.980555Z digest=sha256:8cbb2ccb9bf568ceee5130da9fa6043c2f64f84bcf34d26d8a4b9554aef820f6

Observation f6e5d777-157d-4421-9480-31be1a41639a · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.087178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.087178Z digest=sha256:ed60c09664803cba1c6abfb03673f4f37dc93aa237e3238cba2ecc3c94527668

Observation a9a1aa62-1b7b-41d2-9abd-0de5e46b7720 · outbound

This paper cites Exploring video quality assessment on user generated contents from aesthetic and technical perspectives.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Exploring video quality assessment on user generated contents from aesthetic and technical perspectives

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:09.038369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:07.174284Z digest=sha256:afdb36c5ef24893ac33e531c35462b7d43017002641685d69a9e72a0783d9f96

Observation 7cf0cde8-aba3-4dc7-b04d-146c968e51e1 · outbound

This paper cites Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Tune-a-video: One-shot tun- ing of image diffusion models for text-to-video generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.865062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:07.254277Z digest=sha256:564afcae8c480a1834abbf78820b070e9fa934bbb169a634cf1fa49a5e603493

Observation ffa0efa5-6e01-4291-81c3-745e8b7e73d2 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Rerender a video: Zero-shot text-guided video-to-video translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.669898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:07.339187Z digest=sha256:96986dd9866df131829017d7f34da14d44a31cdff9f8d9356205ed998612ef92

Observation a13928cf-a1bb-4530-9664-af5f101d3778 · outbound

This paper cites Fresco: Spatial-temporal correspondence for zero-shot video translation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis Fresco: Spatial-temporal correspondence for zero-shot video translation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:24:08.509404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:07.427186Z digest=sha256:9892f4a421a8d598ea5ca212e0647a2440cf0907c808c78971408f8f950632da

Observation 60f5666e-ec18-497c-99ac-9f1789767b36 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.516251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.516251Z digest=sha256:18b76f0ddfded0c5c7a1bc77f704633b894daf60fc38d1b14429fb315ccb856b

Observation 97fe92c6-dd19-48d1-8024-14ca4544c834 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.602629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.602629Z digest=sha256:6a72aba75de61af276a790a176a7712b2f6bcdada0f7b7d259dca039df4aea8c

Observation 6f4480d7-b3f2-4914-bd8d-0eb6e57c9371 · outbound

This paper cites ControlVideo: Training-free Controllable Text-to-Video Generation.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis ControlVideo: Training-free Controllable Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:07.699049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:07.699049Z digest=sha256:c29f129282d4ec25613bf551d035e4b3ad43b4ae04ba8e37a0ce5da175e9c36f

Observation ec343ff1-a59f-4dbf-ac32-4a867b6a9583 · outbound

This paper cites VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models.

Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video Synthesis VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:24:08.016744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:24:07.779000Z digest=sha256:93dca838b48e495bd64f7ffd59bd091a31c997ceaf1d6e14b6c6c6dff4b5de54

Pith citing papers

No inbound Pith citation observations are available.