REVIEW 7 cited by
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this work, we design a novel two-stage training scheme that can utilize easily obtained datasets (i.e.,image pose pair and pose-free video) and the pre-trained text-to-image (T2I) model to obtain the pose-controllable character videos. Specifically, in the first stage, only the keypoint-image pairs are used only for a controllable text-to-image generation. We learn a zero-initialized convolutional encoder to encode the pose information. In the second stage, we finetune the motion of the above network via a pose-free video dataset by adding the learnable temporal self-attention and reformed cross-frame self-attention blocks. Powered by our new designs, our method successfully generates continuously pose-controllable character videos while keeps the editing and concept composition ability of the pre-trained T2I model. The code and models will be made publicly available.
Forward citations
Cited by 7 Pith papers
-
MultiAnimate: A Unified Framework for Controllable Multi-Character Animation
A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.
-
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
A feed-forward transformer model that adds SMPL-X neural-texture pose conditioning to LVSM, enabling single-pass human novel-view and novel-pose synthesis that surpasses prior generalizable methods on four benchmarks.
-
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
The paper maps how storytellers prefer to use AI-generated text, audio, images, videos, and 3D content to augment AR stories, based on a 223-video analysis and two user studies with 30 participants.
-
Wan-Animate-2: Pushing the Application Boundaries of Character Animation
An end-to-end Diffusion Transformer animates a still character from a driving video without motion extractors, with optional text camera control and a real-time streaming variant.
-
Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization
HPC-ColPali compresses ColPali's patch embeddings via K-means quantization, attention-based pruning, and binary Hamming search, but all reported numbers are estimates, not measurements.
Discussion (0). Continue with ORCID to comment.