Pith. sign in

REVIEW 7 cited by

Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01186 v2 pith:NAO3QES5 submitted 2023-04-03 cs.CV

classification cs.CV
keywords videoscharacterposepose-controllablepose-freedatasetgenerationmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating text-editable and pose-controllable character videos have an imperious demand in creating various digital human. Nevertheless, this task has been restricted by the absence of a comprehensive dataset featuring paired video-pose captions and the generative prior models for videos. In this work, we design a novel two-stage training scheme that can utilize easily obtained datasets (i.e.,image pose pair and pose-free video) and the pre-trained text-to-image (T2I) model to obtain the pose-controllable character videos. Specifically, in the first stage, only the keypoint-image pairs are used only for a controllable text-to-image generation. We learn a zero-initialized convolutional encoder to encode the pose information. In the second stage, we finetune the motion of the above network via a pose-free video dataset by adding the learnable temporal self-attention and reformed cross-frame self-attention blocks. Powered by our new designs, our method successfully generates continuously pose-controllable character videos while keeps the editing and concept composition ability of the pre-trained T2I model. The code and models will be made publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiAnimate: A Unified Framework for Controllable Multi-Character Animation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A diffusion-based framework that animates multiple characters in one scene from separate reference images and pose sequences while preserving each character's identity.

  2. 3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.

  3. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0 of 10

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  4. HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A feed-forward transformer model that adds SMPL-X neural-texture pose conditioning to LVSM, enabling single-pass human novel-view and novel-pose synthesis that surpasses prior generalizable methods on four benchmarks.

  5. An Exploratory Study on Multi-modal Generative AI in AR Storytelling

    cs.HC 2025-05 conditional novelty 6.0 of 10

    The paper maps how storytellers prefer to use AI-generated text, audio, images, videos, and 3D content to augment AR stories, based on a 223-video analysis and two user studies with 30 participants.

  6. Wan-Animate-2: Pushing the Application Boundaries of Character Animation

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An end-to-end Diffusion Transformer animates a still character from a driving video without motion extractors, with optional text camera control and a real-time streaming variant.

  7. Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization

    cs.IR 2025-06 reject novelty 4.0 of 10

    HPC-ColPali compresses ColPali's patch embeddings via K-means quantization, attention-based pruning, and binary Hamming search, but all reported numbers are estimates, not measurements.

Pith tools