GenHSI is a training-free three-stage pipeline that turns a scene image, character image, and complex HSI prompt into long videos with plausible chained interactions by generating atomic actions, 3D keyframes via 2D inpainting plus optimization, and then feeding them to pre-trained video diffusion.
Magic insert: Style-aware drag-and-drop
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2representative citing papers
Raw CSD cosine is not a calibrated absolute style score for many artists; a discrimination-gap diagnostic flags the failures and CSLS on the frozen backbone corrects most of them.
citing papers explorer
-
GenHSI: Controllable Generation of Human-Scene Interaction Videos
GenHSI is a training-free three-stage pipeline that turns a scene image, character image, and complex HSI prompt into long videos with plausible chained interactions by generating atomic actions, 3D keyframes via 2D inpainting plus optimization, and then feeding them to pre-trained video diffusion.
-
When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
Raw CSD cosine is not a calibrated absolute style score for many artists; a discrimination-gap diagnostic flags the failures and CSLS on the frozen backbone corrects most of them.