Pith. sign in

Paper Citation Record · LEDGER

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

As of 16 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 3 inbound Pith citation observations for arXiv:2411.16199.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16199 v6

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:29:59.155701Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:59:07.294813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:19:57.945426Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a209f8c4-9825-4c3b-83d0-72712a917939 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.911510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.911510Z digest=sha256:f4cfc21c48223f16cad0987b8d5d6527f2a144d3f2ac799dba62ee78d9e789d5

Observation ae78e377-e94c-434c-beb5-daf85b20551d · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, 2021.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Is space-time attention all you need for video understanding? In ICML, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.901470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.917101Z digest=sha256:52565a0645811fa9655b33b3c34e1806df3dfd68c039184be22d112f60fe9a40

Observation 0b09ca38-e901-43f5-b616-2b975ff4c6a7 · outbound

This paper cites Sigmoid- weighted linear units for neural network function approxi- mation in reinforcement learning.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Sigmoid- weighted linear units for neural network function approxi- mation in reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.922100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.922100Z digest=sha256:1f6bfee75853068894fa366a4885a4532484335a0f4b4afda6642027feff1d38

Observation 494e5c2e-5ede-4ef7-a7e3-637a837b1294 · outbound

This paper cites Digital image process- ing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Digital image process- ing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.874984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.927030Z digest=sha256:e7eab51458d3ff1d632541b9265564a5d00b2ccafbb2dcd317b40448280d2bb6

Observation 9f3a6616-4f9c-4b07-8000-43dc8dc4df83 · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.932800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.932800Z digest=sha256:176c0dc2b9e51fd326df9e784d1ad6219bc80c914477dfa811bcc0b5fb609f1b

Observation 0a40730e-ceae-43eb-a96e-c451c510611e · outbound

This paper cites Denoising diffu- sion probabilistic models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Denoising diffu- sion probabilistic models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.938005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.938005Z digest=sha256:5c31d2bc5ab9df3c89890074fc75bb4779c1cb50b4da65bee797ddb128133401

Observation 178c04bb-a860-43a0-a855-f36f4677cf5f · outbound

This paper cites Video dif- fusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Video dif- fusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.942819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.942819Z digest=sha256:003b54418091c0f69adf047f6e905ed015b4f0a06a3266777501c023af7eab3f

Observation 6d075ec8-f105-43ac-ba66-a3146a61a6fc · outbound

This paper cites L-DiffER: Single image reflection removal with language-based diffusion model.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation L-DiffER: Single image reflection removal with language-based diffusion model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.840348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.947860Z digest=sha256:14938b924240882b212f2aa9373c9d8348e6531960cc4c630bae192346bb2732

Observation a798eb74-3632-4963-b910-3a685e1f6701 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.952504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.952504Z digest=sha256:ff88a86e65493fddca8a072f42a1b442188f9ecf2f5c3ece94a2d9a2cd44a0c1

Observation 57f24a54-5572-40a2-a64d-0bce29ecc929 · outbound

This paper cites VBench: Comprehensive bench- mark suite for video generative models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VBench: Comprehensive bench- mark suite for video generative models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.815586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.957529Z digest=sha256:ac1b9980acb996ad120f97fac03cded7a350661ab913a4c7af650aef79a7c5f4

Observation 635f8577-bc7d-4001-a36b-84588ba375d9 · outbound

This paper cites Scope of va- lidity of psnr in image/video quality assessment.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Scope of va- lidity of psnr in image/video quality assessment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.962036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.962036Z digest=sha256:0f78c58ec573e48ac934bb1a19a1a56f6626f6a2684aa4c993892d0f2f040259

Observation d8f5a263-d1fe-40c8-8fa7-71be048b3da6 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Perceptual losses for real-time style transfer and super-resolution

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.790511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.966629Z digest=sha256:2a19b9910c4057b77c608de3bd575697be12c1308d873e28c0a539925b6744e3

Observation c704d836-70c9-4185-9e13-850e183a5a87 · outbound

This paper cites RA VE: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation RA VE: Randomized noise shuf- fling for fast and consistent video editing with diffusion mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.775947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.971285Z digest=sha256:7c0a8a1d36712e8d1694977f89208e0bf438f31e7e96afcea6cc08676876fa2d

Observation 00619ca1-e25f-46fb-8e76-8afc95c88526 · outbound

This paper cites Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Text2Video-Zero: Text- to-image diffusion models are zero-shot video generators

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.761493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.975867Z digest=sha256:aaf01933a860a730cee3f738813af627c61c8ebc7ae9d656b986d563bb90ccf9

Observation f1fffb00-f884-4a86-8157-45787a27e660 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.980712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.980712Z digest=sha256:30db03acceb78c7d6c997dd19c58c81207a13516c76ddb722420267512e286d9

Observation 2d485f0c-3a39-4f67-bd6f-61225de5ed1c · outbound

This paper cites On information and sufficiency.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation On information and sufficiency

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.746579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.985587Z digest=sha256:503d4fc1a350531dedac6b782ea35c7af2f5247b09253e28b253d24b097aaafa

Observation 3a21e06c-be5a-4c50-8d48-94813532d102 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Multi-concept customization of text-to-image diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:58.990083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:58.990083Z digest=sha256:2cf7c16c483bb918ce813ec555adad64a6d97b073ea9b583a60a71dc84f16b39

Observation 6644b61f-a134-472a-950f-518710e9d908 · outbound

This paper cites VidToMe: Video token merging for zero-shot video editing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VidToMe: Video token merging for zero-shot video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.721303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.994700Z digest=sha256:7a8ab5aa7757fe1a1829c1fa8e818ba72d3ca8b9eca6c3bad69839c9a9690458

Observation 091f2d57-f861-458b-acb3-9c80d247feaf · outbound

This paper cites Flow matching for generative modeling.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Flow matching for generative modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.703888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:58.999531Z digest=sha256:732d93665a488ebeb7bd8bd931eda419b43fce794592f7654d93e77a83a3ddb6

Observation 58b5fe4d-2941-47bc-9312-bc31b852f894 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.004337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.004337Z digest=sha256:03da5beb23f1c8c074558628a973edce57ffe11c1850e13524002239ea8da6bd

Observation 11e09c46-724f-440d-b156-f34f4c23f5eb · outbound

This paper cites SDEdit: Guided image synthesis and editing with stochastic differential equa- tions.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SDEdit: Guided image synthesis and editing with stochastic differential equa- tions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.688935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.009489Z digest=sha256:3049a7df0b080902f3e34e6946f221d3983eeec5cc72530b6d2cf427343ed3b3

Observation 931ab61b-b006-49eb-a666-b476004b2bde · outbound

This paper cites T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation T2I-Adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.673805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.014192Z digest=sha256:5c31cb10d157a9f8b139f01dd0051a320b06c88c581a09e64a17af57706a6a2e

Observation 87700ab4-1209-4499-8bd7-93eec6ac1b1c · outbound

This paper cites Semantic image synthesis with spatially-adaptive nor- malization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Semantic image synthesis with spatially-adaptive nor- malization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.018700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.018700Z digest=sha256:856053720839158127db4108a7af6fcb6fe0e574bdbfe7121b41d7ee1207d6b2

Observation ccf6aea1-26c5-4d33-81cc-7e93da5d976d · outbound

This paper cites Zero-shot image-to-image translation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Zero-shot image-to-image translation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.648477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.023360Z digest=sha256:a1ec452601f64f635b73bf1674976a307f163a45acc562dcaea36d98938906e3

Observation e653f0fd-d31f-4a94-9769-1cef8a0f4c3f · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation A benchmark dataset and evaluation methodology for video object segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.632660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.027762Z digest=sha256:9572515f1f11b5ebe2751c31a9bb1f8a32ee361c371c7d42fe97d9b66e53330e

Observation be0fbd8b-5578-40b0-9ed8-722974215a48 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation FiLM: Visual reasoning with a general conditioning layer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.616839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.032279Z digest=sha256:2df8289aaa07599e72b06dfe43376cd0396e454a0535a7c4275f48a03659d8fa

Observation b333c489-7d86-426a-b3cb-fa842f4e198d · outbound

This paper cites FateZero: Fus- ing attentions for zero-shot text-based video editing.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation FateZero: Fus- ing attentions for zero-shot text-based video editing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.036622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.036622Z digest=sha256:53965eabd7a8cd59566d7f7124316036ffbc7b2876fd82fd8b1e5436a47b1f07

Observation 325970f3-43b1-40a3-8f34-7724cbb41e0c · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Learn- ing transferable visual models from natural language super- vision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.041135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.041135Z digest=sha256:cf4f9ebd407c1f1c1f75b5ac85dfbbb15150505657cbb2613534684d5e845326

Observation 2a382dc7-8e64-4e0c-9571-b6e4df2c5040 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.581070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.045614Z digest=sha256:5ececb91413f969b46de2b35dc6fdb8b4379f039dcb5209df2e7bd923aef14e5

Observation 62824420-8e27-427f-bc88-8b91bd15877e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SAM 2: Segment Anything in Images and Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.050115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.050115Z digest=sha256:6106f33698583ce129dddae395b80c0e5fbac65edd4495c3818f641fa7118224

Observation 4827f5de-3482-4538-8343-de52729a7769 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation High-resolution image syn- thesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.565128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.055011Z digest=sha256:b913676b86eae90f55f54f0359ab6566c0e4868f067af056ca65d6c7129a1096

Observation fe9b9e2e-e552-40db-a10b-5f5926499019 · outbound

This paper cites DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.059206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.059206Z digest=sha256:342f151ff815684a23cee49435878ee1d59501f1a7d96cb95ef6f099e0ef4b68

Observation a3c21f0d-9d76-4932-be10-eea45f7ad3a1 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.063533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.063533Z digest=sha256:bd59f02b2241d3a66a00ed60233be1c14daa80cad60bf4951739def213991a73

Observation 5aeff902-3a33-48ba-8bca-4dd231792ed2 · outbound

This paper cites Denois- ing diffusion implicit models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Denois- ing diffusion implicit models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.068278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.068278Z digest=sha256:d46e59b24de9d4b51bad449b199c41be01fa289245e12df9bb34ccaef9a5be4d

Observation 132111fa-c026-4746-b7f2-850718a67ec1 · outbound

This paper cites Interpretable 3D human action analysis with temporal convolutional networks.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Interpretable 3D human action analysis with temporal convolutional networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.517401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.072649Z digest=sha256:de48ce8f9ca347a6e15d3c86fa880f4a8fe67dd157fc66f66e2042cae8ef3867

Observation 73fdab96-0d2b-4e16-bc40-035b3de460fb · outbound

This paper cites Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.077307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.077307Z digest=sha256:b3a6396b7420034378d66e4c02967a4f94d8f2ce4a845df5248b212ebb4e1037

Observation 9cfa4a79-8cc5-4f7c-b69e-5a984caa60a7 · outbound

This paper cites VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.081986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.081986Z digest=sha256:39fd30b0bff72b82e981676bc229abc87038e9fac0acc049ac1cd485c62dbc46

Observation 911f6c40-cf1c-4343-9081-4b5b13c862f8 · outbound

This paper cites VideoComposer: Compositional video synthesis with motion controllability.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VideoComposer: Compositional video synthesis with motion controllability

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.502289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.086883Z digest=sha256:82802bf483b60999a5c69414db5495a721a760a0290b2b7fa00085f2affe2b14

Observation 7f0dedbd-07e4-40a2-9b74-5081dafae0b0 · outbound

This paper cites SinSR: diffusion-based image super- resolution in a single step.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation SinSR: diffusion-based image super- resolution in a single step

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.486800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.091254Z digest=sha256:1d5b0395fc5a6a8c0837d678e371fbd3cbdf6fa98200292a68a25aa1aacc3934

Observation 3b1d200e-6315-4006-9434-682eea499c0f · outbound

This paper cites Bovik, Hamid R.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Bovik, Hamid R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.471575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.095667Z digest=sha256:5564eae4ee2847b020366f15c1e00d47d46e2cc6b6d7c3e0b423792b04bd45e6

Observation 2a4bbe82-c3de-4761-9f2f-ed2a1c06299f · outbound

This paper cites L-CAD: Language-based colorization with any-level descriptions using diffusion priors.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation L-CAD: Language-based colorization with any-level descriptions using diffusion priors

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.455860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.100484Z digest=sha256:0fd2a9e083cc58d8795b9789e703b3d45f479e63da0fc631f12d14aca95502cf

Observation f12d2e24-6d17-40a8-9d38-f924e18a5c21 · outbound

This paper cites Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Tune-A-Video: One-shot tuning of image diffusion models for text-to-video generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.104964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.104964Z digest=sha256:96e74c7b027e3eb206bd5858cfcbfd432bf609badf84cd46a2ea3057dc8bfcd7

Observation 0d0c45e4-a981-4e89-a2c9-b6b1db1ebb58 · outbound

This paper cites Group normalization.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Group normalization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.109283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.109283Z digest=sha256:7de044bdbdd9edec01e6737979bd9fa1e05649b6fcf59591a498fc23dc942067

Observation dde3be38-f9f9-48d9-b02e-69ab4951e81d · outbound

This paper cites Holistically-nested edge de- tection.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Holistically-nested edge de- tection

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.421589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.114189Z digest=sha256:c4b230e0555c047032a773d7513f1c2181b553811e84e3ad555bbe371b939e48

Observation 174f776a-ec3d-489d-9cf7-b75536be5995 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.118569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.118569Z digest=sha256:b6ff6679fc4221388c7d2d3e79fb48012a92402aeed020735ab94e5ad964e534

Observation 79054d7b-5931-4890-9dec-e93b9c8041d6 · outbound

This paper cites Depth Anything V2.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Depth Anything V2

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.123262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.123262Z digest=sha256:320d8ea5853783ed772e4fdacaaddf7ea88d41326b8bf92c590014e260c9ba2d

Observation 973855dc-9e1e-4b71-b74c-5252c7268817 · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Rerender a video: Zero-shot text-guided video-to-video translation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.405529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.127893Z digest=sha256:db720371baee0b2bb46357506408fc85c13aa2154eae02bb9704af12beaed8a3

Observation 14fa1c86-405a-4889-8316-3dd6d9fd47fe · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.132657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.132657Z digest=sha256:b3334380fd5e67b37203b5e9caf902c8ac575e8a350dab2b98e214ccec4845fb

Observation f50df8fd-ff5a-4334-b82a-dc442d884ce8 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Language model beats diffusion-tokenizer is key to visual generation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.390205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.137522Z digest=sha256:5e1d07a5d6f3b5eee887477540610d9f37e41f8f7288ba252529c0899ec1ba8f

Observation dba27e5a-846e-4561-b875-75c7d901dc82 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.142094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.142094Z digest=sha256:e424801a7655f4e27f2ca603fcc7703f3aa3816c269d24a1bb23cf9ba8ff28d0

Observation a55b4322-6ddc-4579-b65a-632d6573d69b · outbound

This paper cites ControlVideo: Training-free controllable text-to-video generation.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation ControlVideo: Training-free controllable text-to-video generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.364635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.146524Z digest=sha256:b8639e9c92f8d46951b2ae0c158bbe803012a9abceaef54c4d2423e0cfcc024c

Observation e0f3f341-0811-4925-9b1a-9a9e61bd6291 · outbound

This paper cites Open-Sora: Democratizing efficient video production for all, 2024.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Open-Sora: Democratizing efficient video production for all, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:29:59.349056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T13:29:59.151647Z digest=sha256:a05e9dde3b1081d93140bd1644e65daa2b62e1e993afa13472c7dc1eac359dae

Observation 5749ee67-3660-4968-aec5-4d747236948e · outbound

This paper cites Allegro: Open the Black Box of Commercial-Level Video Generation Model.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation Allegro: Open the Black Box of Commercial-Level Video Generation Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.155701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.155701Z digest=sha256:ffdefff6e963083dcee24985fc0785201a9cca81a38ad090f70014f05589047e

Pith citing papers

Observation 89ee7eed-8e20-4c77-99e1-711171968f18 · inbound

Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion cites this paper.

Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:18:19.535839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T19:14:15.240218Z digest=sha256:890870179a6da76166de1bd901cf47d00fbb31884d3bdb004ca2db1206b4f24c

Observation 69d6b6b6-a0a4-4dd1-8ca9-7aa53a39df81 · inbound

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration cites this paper.

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.947045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T00:45:05.681548Z digest=sha256:fe6a181f1d1584567afc58a155c686f5c5735ab156fb353d33ebcbc3cea6fcb7

Observation b6e8d568-da58-418c-bb76-ff9bc6ee697c · inbound

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration cites this paper.

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration VIRES: Video Instance Repainting via Sketch and Text Guided Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.630129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-01T06:59:07.294813Z digest=sha256:47ca779f5cfbb44d1dc6226cbb85d0db9a60630a505e421317868e14567402de