Pith. sign in

Paper Citation Record · LEDGER

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

As of 21 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2412.10533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10533 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:56:59.884771Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:59:49.398271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T08:02:10.321954Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59bbd1ae-a2aa-4e20-a7d3-d2373a329493 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.708772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.708772Z digest=sha256:ed6ba0206027f8dab9b1b1ee0a4dfc35b269547795b4d1e89b8435d6d7cf3a15

Observation 2da0f429-201e-4bbc-b52f-3584b8011976 · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.712836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.712836Z digest=sha256:a912cf003ff6d6c65ba5252778b748d8a7a448b71c497792bb530748a9f70ce1

Observation 47ccf589-3e8c-4a81-a250-70e485f63dcb · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.716910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.716910Z digest=sha256:713bb6c0caea413f116c76aa8d32ccf4ecf0aef89f7985a6821a3b332c5d70c2

Observation a36975d0-24fe-4705-82b5-f943b630e39b · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.386952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.721156Z digest=sha256:e12a1e47eab7f1c19bae1a350b90dd5f09d74f27bcf60dc322656c84be7b5df0

Observation 39352f1e-6020-4f50-93e1-e8d983bde9e4 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.724691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.724691Z digest=sha256:e2d889cd679c3fb3408fb2be3fc1dc002696c33f5dcbab12d5d2f88a08b943c1

Observation 0c507006-2423-4db7-a323-9863727f7743 · outbound

This paper cites Video generation models as world simulators,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Video generation models as world simulators,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.728227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.728227Z digest=sha256:f1bc6b997975319ef06982ebdc07f02ef4334c3addf383faade5b91510681765

Observation 19a02f93-75f1-43c3-a073-e295672d3991 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Emerg- ing properties in self-supervised vision transformers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.363853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.731701Z digest=sha256:a409182d3b4294d8f21eeaa58e2b18683dc56f8c4513318abd58ae82841fe803

Observation b90e1452-3a61-48ec-8aca-a7dcaaf71de6 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Panda-70m: Captioning 70m videos with multiple cross-modality teachers, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.352581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.735334Z digest=sha256:5dd27a0b4418d78f020e8f99e102296de38ba7ef0fb404cce080aa52f833f8de

Observation 1ad3085a-95b8-41f1-b6ef-dbf825b5011b · outbound

This paper cites Subject-driven Text-to-Image Generation via Apprenticeship Learning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Subject-driven Text-to-Image Generation via Apprenticeship Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.739929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.739929Z digest=sha256:634718b0bf891fe41b3fceffb06a637288ca3509f3cf67aa382e34d40e392834

Observation de93377e-56ca-4595-afec-b447c7392da1 · outbound

This paper cites AnyDoor: Zero-shot Object-level Image Customization.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner AnyDoor: Zero-shot Object-level Image Customization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.743465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.743465Z digest=sha256:a94bb963fff9dead87e445a86d24d47637637403ff2f2c588dd61eb558c06b8b

Observation 98d16545-fab3-48ce-9904-a5a0de3c467c · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.747211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.747211Z digest=sha256:e9d27734fe07781bf15b0b18274e8ff6326637a1318f6e4e0b95f70ffe110ba8

Observation b9daad35-7f56-459e-83d2-4af4696b8dda · outbound

This paper cites An image is worth one word: Personalizing text-to-image gener- ation using textual inversion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner An image is worth one word: Personalizing text-to-image gener- ation using textual inversion

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.334812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.750662Z digest=sha256:a0356c38e939417059bc645069d2602eae1c822a2c66f5c7d4515b7d4c256649

Observation 1c6d2022-1679-4b6c-ad95-00f8dd4db59d · outbound

This paper cites Photorealistic video generation with diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Photorealistic video generation with diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.323368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.753971Z digest=sha256:3e86574cc5fc1bd65427e32d2ba3056b455a138a8949a72e3cc7b870d17716b2

Observation 1a2aa6c9-2160-4eae-9379-631fb3845d01 · outbound

This paper cites Classifier-free diffusion guidance.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Classifier-free diffusion guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.757231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.757231Z digest=sha256:55502075ac5a18cc669ad8534bcb1146c2d08b7b8886b2ac464a682b8ea228fb

Observation 87563d6d-011a-4ccf-b66f-8fceb9f4dd53 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Imagen Video: High Definition Video Generation with Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.760051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.760051Z digest=sha256:22b640d7dd425a0321cd1db6ec269f8951c4733cfcb1a6795763d72df566c296

Observation c5225c12-bf47-4636-b644-ba3ba1acf683 · outbound

This paper cites Vbench: Com- prehensive benchmark suite for video generative models,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vbench: Com- prehensive benchmark suite for video generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.763446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.763446Z digest=sha256:8ccc24d781d85ce01c29fc2706c5b21c5b42260f238ba4cddc05f88181238ad1

Observation 6448f32a-29d8-44cf-a834-14cb6e323c87 · outbound

This paper cites Videobooth: Diffusion-based video generation with image prompts.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Videobooth: Diffusion-based video generation with image prompts

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.300238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.766888Z digest=sha256:7aa5ae112576ae942576f9d5fa53301c2bb2bbace03860d58aa756372643c961

Observation 9b7c4109-6273-4269-836a-ad3c8baaf761 · outbound

This paper cites Auto-Encoding Variational Bayes.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Auto-Encoding Variational Bayes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.769753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.769753Z digest=sha256:b65fe29994638a22e736050d074d65292040cdd7e22db787f3ce2e6f2dd9db3f

Observation 50558660-b6e4-4882-ba4c-733cb2f67264 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.773238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.773238Z digest=sha256:f0318e4945c8975798d7aa682bff4abc4525792867835bbd0ea01cc2906a8c9e

Observation 633e40d1-d54a-4ff4-9af5-814284aae0de · outbound

This paper cites Multi-Concept Customization of Text-to-Image Diffusion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Multi-Concept Customization of Text-to-Image Diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.776657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.776657Z digest=sha256:ed2a5bf82d7cd7cf6ec036de8bd5aee4dd78f55c3dae6c90d6a89dbc3338009b

Observation 6f922097-c920-4097-bf7d-78921a88ae39 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Multi-concept customization of text-to-image diffusion

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.289833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.780665Z digest=sha256:0b65d5249f5949f795e466d136965e3a7f03ce59a22a76d8712a6b87f41222ec

Observation d11d5300-00fa-4a67-be32-82579b415bdd · outbound

This paper cites Flux, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Flux, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.784156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.784156Z digest=sha256:bb04457b0714b5391d103cfb815f2c93e112e7077b6d109f613ca35e8c6800e8

Observation d95009a3-d759-4854-8e37-0299f77eafaa · outbound

This paper cites Luma dream machine, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Luma dream machine, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.272606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.787688Z digest=sha256:58ce1543693088e459ba7de27149e53777c588597b2988e43591eed93f9c6068

Observation e1e03782-b62d-45ab-b5e6-2e7bb7fc09c9 · outbound

This paper cites Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Subject-Diffusion:Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.791130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.791130Z digest=sha256:15508fe0f041010ad8b7b0cea44012f3119687a72415fe5e230c0cd81e14b9d3

Observation d6e1cda3-4448-4b27-a6da-8eb19ecb3b51 · outbound

This paper cites Dinov2: Learning robust visual features with- out supervision, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dinov2: Learning robust visual features with- out supervision, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.261606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.794848Z digest=sha256:249faf9ba4f8ff591283018acc3129a78cbd8dc7435a050da58ad0cb6ee6f9d5

Observation 4b479e08-728f-4663-b945-eddc8005300b · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.798517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.798517Z digest=sha256:19c1d59d461b6bd65d91412090b92c9a66a5f7ef16018687cc96747eeea91569

Observation 40c450c9-ddfe-44be-afe6-d3d52000a588 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.250736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.802465Z digest=sha256:8519adc1a12541c5be499f855a2770533e72ba6e2016ddbe11be2d5778b01ab7

Observation 5e08c81f-3c88-4e61-860c-d37d401936bd · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Movie Gen: A Cast of Media Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.806149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.806149Z digest=sha256:841723e3e3fd97c8d0607d8bf3ea9d518f0c906ac7b91588ffcae3a4243be7e7

Observation 158342f1-584f-4a33-843f-9cb45ecf18c3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Learning transferable visual models from natural language supervi- sion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.240382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.810182Z digest=sha256:3828b32a97a45b49308972156b11b111c1f1d1730c94418ab9bac62758780d85

Observation b73b69ad-680a-4ed4-9775-650c7037ff28 · outbound

This paper cites Zero-shot text-to-image generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Zero-shot text-to-image generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.230002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.813440Z digest=sha256:80ac4d8cf66156ab3267a3a2b6f72132141c431a9265f0c9de37eefad311546d

Observation 559d896d-e333-4552-ba5c-dea380914e72 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks,.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Grounded sam: Assembling open-world models for diverse visual tasks,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.816893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.816893Z digest=sha256:a82a810dab042c15abd94d7f9734f7a7bb4bbb814e1908833c88d8d6d7dd6b8e

Observation 8a12b6b9-2fdc-411d-8db7-aa32eda2861e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner High-resolution image synthesis with latent diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.820341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.820341Z digest=sha256:dbf61f64fa11d34c865626bb5a621f7d6e2cc5834a79c08fe7cbf3418633603d

Observation a8414ba3-8c42-49b0-b472-3d24bee5d2c6 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.207888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.823699Z digest=sha256:700d152fe3ec91e63e7ebe6e37b75f6b27c356b909f507c228631331e0feae9f

Observation 21ce158a-a063-4d99-b97b-da9aa34c843b · outbound

This paper cites Gen-3 alpha, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Gen-3 alpha, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.198475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.827139Z digest=sha256:a5ec44c701605a2170257c20a5cdb17926e2a0ad134035b8d1880b9b838be222

Observation 5491f2eb-6d4a-4dbb-bd1b-04c6b0d53833 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Photorealistic text-to-image diffusion models with deep language understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.830494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.830494Z digest=sha256:ec5e65b0b1765807130250dcc09a40d131537fc5d15a4d71693c31d85ca46402

Observation db47afd9-34be-4b37-9adb-275cf0cb64e5 · outbound

This paper cites InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.834021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.834021Z digest=sha256:94cef2c8ceda537085c4392de02ca7e04d5a302cab92490f898bc58b06084008

Observation 2571f636-ea87-4d93-a27c-5cc3282ffd47 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Make-a-video: Text-to-video generation without text-video data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.182057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.837693Z digest=sha256:7ea367d6b6dff01355d2c5d5e2bb79ac4080fc14979262d9779bfcc9f1bdadc1

Observation 140dc4fc-bdbf-4cee-8409-0842352a484e · outbound

This paper cites Vidu-1.5, 2024.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Vidu-1.5, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.171942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.841160Z digest=sha256:82839d491c7dbde190d3d1af792c13cee903e1643e75ad4dc8fca3de6b1f6443

Observation 74d9c82e-6976-4601-90e8-36461d81ec3a · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Raft: Recurrent all-pairs field transforms for optical flow

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.844534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.844534Z digest=sha256:d2f01ff873fba4dabff52ee7d1012d0a3250fa04f94b806dce442e6c10c56e41

Observation b31fab30-08df-4b24-a023-d72695fb7a04 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.847951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.847951Z digest=sha256:6b55ae17c5ba06d380929048abe124e50dfac95f607c0cb773f574bfc68b1807

Observation 8cbef5c0-658a-4364-9608-70e79551bd21 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.851385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.851385Z digest=sha256:8e602be5f95ffe076fe00a9faddb42e6e33317092024879a79efade57cf9a6a5

Observation 2773f20a-be5a-40e7-b5f4-45f16055ae43 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.155337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.854770Z digest=sha256:6af7f3e5633f7e44bc362951b040edecbc6bb8c9d97b4bdf101dafb0edd4af58

Observation db47e8a1-1982-4929-bca5-aa81b5b3b573 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.857766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.857766Z digest=sha256:4e046fb0bfae7b7129b0e384c18a0247e5e7b8bdf2ba5e8f7ea7b3c5f91dca34

Observation aca6d5a6-5ab2-4d32-9675-11172a53646f · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Dreamvideo: Composing your dream videos with customized subject and motion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.144278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.861277Z digest=sha256:d85b63012dd334a5c6d0bfef33534bdaafa70ec05dc932f2dd6703e6f137ed38

Observation c9b27005-25cd-49eb-8902-43ac6720a73d · outbound

This paper cites DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.864206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.864206Z digest=sha256:0e0f7b76e363bc380e3cf0f5dcc094e0ad8279b21c91881f96bbc17e6d81f4ba

Observation 5621611e-500a-407e-ae53-d9bc6eaf7c1a · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.867818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.867818Z digest=sha256:44543ec2592baf612926a9adf0cc28b19f3e616eea96b09dab38838f0b231879

Observation e5ab495c-cb78-4a2b-9de3-372923b1b56c · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.870792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.870792Z digest=sha256:d6c4c4b197010dcb17d431e174cdfed37e9010c874a8e99a079a994ca2ffa9d7

Observation 4f4d8f0a-e63a-4cdb-ba46-260d45622ead · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.874312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.874312Z digest=sha256:0387fb0ae92a48f4b437dc1df20383c71c3ea313c4b9192c4fb40a4805e2a61d

Observation 1e2374a1-e3d0-4798-a079-795cecbda668 · outbound

This paper cites Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:56:59.877924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:56:59.877924Z digest=sha256:13171a5f80600f182ba144fcfe28ecae044ba9303c0abd94ff3c87dcb4ea2611

Observation 3a35599c-b16c-48cb-aa01-978009ec18a0 · outbound

This paper cites Cus- tomization assistant for text-to-image generation.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner Cus- tomization assistant for text-to-image generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.127398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.881408Z digest=sha256:f24178919866890d18fafa610d69142cb95ca1bcd41bed0f95f92cdd0542d9ed

Observation 0690ad41-8ad0-4b17-a153-70a55b4b9120 · outbound

This paper cites pencil drawing drawn by a hand.

SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner pencil drawing drawn by a hand

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:57:00.116346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T15:56:59.884771Z digest=sha256:523a94bb5909deeaff13dd059d3e95acaac5d72fe1097097d7b385757acd45c0

Pith citing papers

Observation 4fa29253-1ce0-4f96-9a5e-93e646f1a396 · inbound

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation cites this paper.

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:10.325181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T07:59:49.398271Z digest=sha256:4f997c8ae6f52114365b42c5b699ac24d87660feeccbeb1d43ff889b527e1804