Pith. sign in

Paper Citation Record · LEDGER

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

As of 15 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 3 inbound Pith citation observations for arXiv:2411.16199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16199 v6

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:29:59.155701Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:59:07.294813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:19:57.945426Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a209f8c4-9825-4c3b-83d0-72712a917939 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.911510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.911510Z digest=sha256:f4cfc21c48223f16cad0987b8d5d6527f2a144d3f2ac799dba62ee78d9e789d5

Observation ae78e377-e94c-434c-beb5-daf85b20551d · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Is space-time attention all you need for video understanding? In ICML, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.901470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.917101Z digest=sha256:fb5e54bf7c44bf352cf5e93bfc91a612b26be2b633a22b0ab32129b6fb850fb8

Observation 0b09ca38-e901-43f5-b616-2b975ff4c6a7 · outbound

This paper cites Sigmoid- weighted linear units for neural network function approxi- mation in reinforcement learning.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Sigmoid- weighted linear units for neural network function approxi- mation in reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.922100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.922100Z digest=sha256:1f6bfee75853068894fa366a4885a4532484335a0f4b4afda6642027feff1d38

Observation 494e5c2e-5ede-4ef7-a7e3-637a837b1294 · outbound

This paper cites Digital image process- ing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Digital image process- ing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.874984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.927030Z digest=sha256:a402551d68bc1039182164b413bedd7025a367e879ef2e6e7683bab30e90f16f

Observation 9f3a6616-4f9c-4b07-8000-43dc8dc4df83 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.932800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.932800Z digest=sha256:176c0dc2b9e51fd326df9e784d1ad6219bc80c914477dfa811bcc0b5fb609f1b

Observation 0a40730e-ceae-43eb-a96e-c451c510611e · outbound

This paper cites Denoising diffu- sion probabilistic models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Denoising diffu- sion probabilistic models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.938005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.938005Z digest=sha256:5c31d2bc5ab9df3c89890074fc75bb4779c1cb50b4da65bee797ddb128133401

Observation 178c04bb-a860-43a0-a855-f36f4677cf5f · outbound

This paper cites Video dif- fusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Video dif- fusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.942819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.942819Z digest=sha256:003b54418091c0f69adf047f6e905ed015b4f0a06a3266777501c023af7eab3f

Observation 6d075ec8-f105-43ac-ba66-a3146a61a6fc · outbound

This paper cites L-DiffER: Single image reflection removal with language-based diffusion model.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation L-DiffER: Single image reflection removal with language-based diffusion model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.840348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.947860Z digest=sha256:5629a5b136cc789a88e1cb3a27e8e92e0be912318ace7bc1992fa6e241a3d750

Observation a798eb74-3632-4963-b910-3a685e1f6701 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.952504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.952504Z digest=sha256:ff88a86e65493fddca8a072f42a1b442188f9ecf2f5c3ece94a2d9a2cd44a0c1

Observation 57f24a54-5572-40a2-a64d-0bce29ecc929 · outbound

This paper cites VBench: Comprehensive bench- mark suite for video generative models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VBench: Comprehensive bench- mark suite for video generative models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.815586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.957529Z digest=sha256:7aae383171ae54a990ce26dda2a617de4b1472c6a8a607ffb04dd238bdb457c8

Observation 635f8577-bc7d-4001-a36b-84588ba375d9 · outbound

This paper cites Scope of va- lidity of psnr in image/video quality assessment.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Scope of va- lidity of psnr in image/video quality assessment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.962036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.962036Z digest=sha256:0f78c58ec573e48ac934bb1a19a1a56f6626f6a2684aa4c993892d0f2f040259

Observation d8f5a263-d1fe-40c8-8fa7-71be048b3da6 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Perceptual losses for real-time style transfer and super-resolution

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.790511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.966629Z digest=sha256:10507b6c7972b9976336434aad10a4a1b503ba281d726ab9eb114dba9c92b162

Observation c704d836-70c9-4185-9e13-850e183a5a87 · outbound

This paper cites RA VE: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation RA VE: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.775947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.971285Z digest=sha256:f52fbbd9a4525a3e6747e27f5e660e3f4c798fe344c53553f2c83993a09a7e15

Observation 00619ca1-e25f-46fb-8e76-8afc95c88526 · outbound

This paper cites Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.761493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.975867Z digest=sha256:33c68d86bcc6f61a1ee11a582c308292d23f10712b1c74e7438d4e1bab0778c7

Observation f1fffb00-f884-4a86-8157-45787a27e660 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.980712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.980712Z digest=sha256:30db03acceb78c7d6c997dd19c58c81207a13516c76ddb722420267512e286d9

Observation 2d485f0c-3a39-4f67-bd6f-61225de5ed1c · outbound

This paper cites On information and sufficiency.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation On information and sufficiency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.746579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.985587Z digest=sha256:b27a294dcec1b3bf8d3619933bc662b5acbb35de176cfc64172400a013a99153

Observation 3a21e06c-be5a-4c50-8d48-94813532d102 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Multi-concept customization of text-to-image diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.990083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.990083Z digest=sha256:2cf7c16c483bb918ce813ec555adad64a6d97b073ea9b583a60a71dc84f16b39

Observation 6644b61f-a134-472a-950f-518710e9d908 · outbound

This paper cites VidToMe: Video token merging for zero-shot video editing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VidToMe: Video token merging for zero-shot video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.721303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.994700Z digest=sha256:c71fb84464ac916862162a1c490d3af9a4ece5cd0f1efb0d72aa661efe4528a8

Observation 091f2d57-f861-458b-acb3-9c80d247feaf · outbound

This paper cites Flow matching for generative modeling.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Flow matching for generative modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.703888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:58.999531Z digest=sha256:44333aa1b2ee64b62db8d840bf5ab5c4b1ae76e4d2824b58ac24f4b0dc44e7fe

Observation 58b5fe4d-2941-47bc-9312-bc31b852f894 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.004337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.004337Z digest=sha256:03da5beb23f1c8c074558628a973edce57ffe11c1850e13524002239ea8da6bd

Observation 11e09c46-724f-440d-b156-f34f4c23f5eb · outbound

This paper cites SDEdit: Guided image synthesis and editing with stochastic differential equa- tions.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SDEdit: Guided image synthesis and editing with stochastic differential equa- tions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.688935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.009489Z digest=sha256:c7330c29641f7b375efeb4431c4debbbf95dca38fae3f546a719f3367f991ae7

Observation 931ab61b-b006-49eb-a666-b476004b2bde · outbound

This paper cites T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.673805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.014192Z digest=sha256:56c068027ac6f99d0b87c0773f84b0fe18008016e53a88cf4d0d51f9435e0bae

Observation 87700ab4-1209-4499-8bd7-93eec6ac1b1c · outbound

This paper cites Semantic image synthesis with spatially-adaptive nor- malization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Semantic image synthesis with spatially-adaptive nor- malization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.018700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.018700Z digest=sha256:856053720839158127db4108a7af6fcb6fe0e574bdbfe7121b41d7ee1207d6b2

Observation ccf6aea1-26c5-4d33-81cc-7e93da5d976d · outbound

This paper cites Zero-shot image-to-image translation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Zero-shot image-to-image translation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.648477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.023360Z digest=sha256:65b7cbb1e4ae12ebb879936c4c9bf061de5a51384ad9ea8ae5ac49e185f086ab

Observation e653f0fd-d31f-4a94-9769-1cef8a0f4c3f · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation A benchmark dataset and evaluation methodology for video object segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.632660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.027762Z digest=sha256:d3f06815df06e3fc198dcb23e5a3412597c66e07b66ec57e8bae798533216cab

Observation be0fbd8b-5578-40b0-9ed8-722974215a48 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation FiLM: Visual reasoning with a general conditioning layer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.616839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.032279Z digest=sha256:bdf10c9957f239feff27d4b741cc940b2e9b2557faae904237997271db18ab8d

Observation b333c489-7d86-426a-b3cb-fa842f4e198d · outbound

This paper cites FateZero: Fus- ing attentions for zero-shot text-based video editing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation FateZero: Fus- ing attentions for zero-shot text-based video editing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.036622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.036622Z digest=sha256:53965eabd7a8cd59566d7f7124316036ffbc7b2876fd82fd8b1e5436a47b1f07

Observation 325970f3-43b1-40a3-8f34-7724cbb41e0c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.041135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.041135Z digest=sha256:cf4f9ebd407c1f1c1f75b5ac85dfbbb15150505657cbb2613534684d5e845326

Observation 2a382dc7-8e64-4e0c-9571-b6e4df2c5040 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.581070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.045614Z digest=sha256:3c7a6db2b7f7dd3a08d844a79c4abfd29f3a22dac92a04927155b25d3a49862a

Observation 62824420-8e27-427f-bc88-8b91bd15877e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SAM 2: Segment Anything in Images and Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.050115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.050115Z digest=sha256:6106f33698583ce129dddae395b80c0e5fbac65edd4495c3818f641fa7118224

Observation 4827f5de-3482-4538-8343-de52729a7769 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation High-resolution image syn- thesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.565128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.055011Z digest=sha256:711d3254c9a7de51927130be8085364511c3b8d4ddcace8f9f4066639086bd44

Observation fe9b9e2e-e552-40db-a10b-5f5926499019 · outbound

This paper cites DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.059206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.059206Z digest=sha256:342f151ff815684a23cee49435878ee1d59501f1a7d96cb95ef6f099e0ef4b68

Observation a3c21f0d-9d76-4932-be10-eea45f7ad3a1 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.063533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.063533Z digest=sha256:bd59f02b2241d3a66a00ed60233be1c14daa80cad60bf4951739def213991a73

Observation 5aeff902-3a33-48ba-8bca-4dd231792ed2 · outbound

This paper cites Denois- ing diffusion implicit models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Denois- ing diffusion implicit models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.068278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.068278Z digest=sha256:d46e59b24de9d4b51bad449b199c41be01fa289245e12df9bb34ccaef9a5be4d

Observation 132111fa-c026-4746-b7f2-850718a67ec1 · outbound

This paper cites Interpretable 3D human action analysis with temporal convolutional networks.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Interpretable 3D human action analysis with temporal convolutional networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.517401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.072649Z digest=sha256:8a89075972c377b03b4a1faf80fbd6f2e47fb9f4ed1013b578149bb259978a3c

Observation 73fdab96-0d2b-4e16-bc40-035b3de460fb · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.077307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.077307Z digest=sha256:b3a6396b7420034378d66e4c02967a4f94d8f2ce4a845df5248b212ebb4e1037

Observation 9cfa4a79-8cc5-4f7c-b69e-5a984caa60a7 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.081986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.081986Z digest=sha256:4ff1b62a667bfabbee9b92828fdc48c519decef166712b6ee68bb7ffaafcf99c

Observation 911f6c40-cf1c-4343-9081-4b5b13c862f8 · outbound

This paper cites VideoComposer: Compositional video synthesis with motion controllability.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VideoComposer: Compositional video synthesis with motion controllability

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.502289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.086883Z digest=sha256:2260757cfb6f145db6e66738d033c2a8e8e056f028efcae228d771f124f6dbba

Observation 7f0dedbd-07e4-40a2-9b74-5081dafae0b0 · outbound

This paper cites SinSR: diffusion-based image super- resolution in a single step.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SinSR: diffusion-based image super- resolution in a single step

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.486800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.091254Z digest=sha256:046c39640a0804355d77ac361780437a53a2978d23a4891e4b23c54888971fe3

Observation 3b1d200e-6315-4006-9434-682eea499c0f · outbound

This paper cites Bovik, Hamid R.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Bovik, Hamid R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.471575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.095667Z digest=sha256:707665090bd9d23de656137317dc8c3b2bbf028f6e5b03b46f076f43a51d1cf2

Observation 2a4bbe82-c3de-4761-9f2f-ed2a1c06299f · outbound

This paper cites L-CAD: Language-based colorization with any-level descriptions using diffusion priors.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation L-CAD: Language-based colorization with any-level descriptions using diffusion priors

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.455860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.100484Z digest=sha256:584d28b8a7fe59d32eab52aae46e742b6a80a952845ed8e96189c4e98a1b564f

Observation f12d2e24-6d17-40a8-9d38-f924e18a5c21 · outbound

This paper cites Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.104964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.104964Z digest=sha256:96e74c7b027e3eb206bd5858cfcbfd432bf609badf84cd46a2ea3057dc8bfcd7

Observation 0d0c45e4-a981-4e89-a2c9-b6b1db1ebb58 · outbound

This paper cites Group normalization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Group normalization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.109283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.109283Z digest=sha256:7de044bdbdd9edec01e6737979bd9fa1e05649b6fcf59591a498fc23dc942067

Observation dde3be38-f9f9-48d9-b02e-69ab4951e81d · outbound

This paper cites Holistically-nested edge de- tection.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Holistically-nested edge de- tection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.421589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.114189Z digest=sha256:ca85b3a3db1372c97bb95e76c86c9373c2d99f0a9bbe24fcd38c2f8c7f2a0101

Observation 174f776a-ec3d-489d-9cf7-b75536be5995 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.118569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.118569Z digest=sha256:b6ff6679fc4221388c7d2d3e79fb48012a92402aeed020735ab94e5ad964e534

Observation 79054d7b-5931-4890-9dec-e93b9c8041d6 · outbound

This paper cites Depth Anything V2.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Depth Anything V2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.123262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.123262Z digest=sha256:320d8ea5853783ed772e4fdacaaddf7ea88d41326b8bf92c590014e260c9ba2d

Observation 973855dc-9e1e-4b71-b74c-5252c7268817 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Rerender a video: Zero-shot text-guided video-to-video translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.405529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.127893Z digest=sha256:95187dd9a504c8db0a29bb007226e405415f260924db6cf1d5735acce6b8ede4

Observation 14fa1c86-405a-4889-8316-3dd6d9fd47fe · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.132657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.132657Z digest=sha256:5df3f90c4a17d9c25092c3aeb690082ab7ef6c9108a9587887ddb2bca5000fa8

Observation f50df8fd-ff5a-4334-b82a-dc442d884ce8 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Language model beats diffusion-tokenizer is key to visual generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.390205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.137522Z digest=sha256:3a8616d85d5efdbd5bbeedbc157160413d8f121b7285dd6eedaa600810ee2e66

Observation dba27e5a-846e-4561-b875-75c7d901dc82 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.142094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.142094Z digest=sha256:e424801a7655f4e27f2ca603fcc7703f3aa3816c269d24a1bb23cf9ba8ff28d0

Observation a55b4322-6ddc-4579-b65a-632d6573d69b · outbound

This paper cites ControlVideo: Training-free controllable text-to-video generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation ControlVideo: Training-free controllable text-to-video generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.364635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.146524Z digest=sha256:833aabe8440f607bcd5b86822411ab0a2751cb14c197e1ab1be8884dcccbdcad

Observation e0f3f341-0811-4925-9b1a-9a9e61bd6291 · outbound

This paper cites Open-Sora: Democratizing efficient video production for all, 2024.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Open-Sora: Democratizing efficient video production for all, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.349056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T13:29:59.151647Z digest=sha256:c7f00657ff3e7ffd84e5049d137f05e0808776e34a76ea27492203cad58979d0

Observation 5749ee67-3660-4968-aec5-4d747236948e · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.155701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.155701Z digest=sha256:ffdefff6e963083dcee24985fc0785201a9cca81a38ad090f70014f05589047e

Pith citing papers

Observation 89ee7eed-8e20-4c77-99e1-711171968f18 · inbound

Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion cites this paper.

Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:18:19.535839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T19:14:15.240218Z digest=sha256:52f84ef884a478c0fd824a348fe8f976ea7bb489e51025ccffa9822d9bc369a3

Observation 69d6b6b6-a0a4-4dd1-8ca9-7aa53a39df81 · inbound

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration cites this paper.

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.947045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T00:45:05.681548Z digest=sha256:76943b6c7145f3baee32065f02055c2e2b99525c29c44b9f9b3a0b0bd4494e04

Observation b6e8d568-da58-418c-bb76-ff9bc6ee697c · inbound

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration cites this paper.

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.630129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T06:59:07.294813Z digest=sha256:c3793e992752f84aa5bd880372d4caf50cba2309d5ccaf1cba3a91eab76902ab