Pith. sign in

REVIEW 19 cited by

CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10592 v1 pith:N33L5QDY submitted 2025-03-13 cs.CV

classification cs.CV
keywords videodynamiccameraexplorationscenecamera-controlledcameractrldynamics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics and limited range of viewpoints when generating videos with large camera movement. We take an approach that progressively expands the generation of dynamic scenes -- first enhancing dynamic content within individual video clip, then extending this capability to create seamless explorations across broad viewpoint ranges. Specifically, we construct a dataset featuring a large degree of dynamics with camera parameter annotations for training while designing a lightweight camera injection module and training scheme to preserve dynamics of the pretrained models. Building on these improved single-clip techniques, we enable extended scene exploration by allowing users to iteratively specify camera trajectories for generating coherent video sequences. Experiments across diverse scenarios demonstrate that CameraCtrl Ii enables camera-controlled dynamic scene synthesis with substantially wider spatial exploration than previous approaches.

Discussion (0). Sign in to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rays as Pixels: Learning A Joint Distribution of Videos and Camera Trajectories

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Encoding cameras as pixel-aligned raxels lets one video diffusion model jointly denoise video and trajectories, supporting pose estimation, controlled generation, and joint synthesis.

  2. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation

    cs.CV 2025-12 conditional novelty 7.0 of 10

    GimbalDiffusion lets text-to-video models follow absolute, gravity-aligned camera rotations by training on random crops from 360° video with forward-facing captions.

  3. ID-V2V: Identity-Preserving Video Restylization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ID-V2V restyles video by conditioning a diffusion model on edited keyframes, depth, relit faces, and face normals, so scene edits propagate while facial identity and performance are preserved.

  4. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

    cs.CV 2026-04 conditional novelty 6.0 of 10

    SymphoMotion jointly controls camera trajectories and depth-aware object dynamics inside one video diffusion model, supported by the new RealCOD-25K real-world paired-motion dataset.

  5. GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...

  6. UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A video-generation world model that warps positional encodings of memory frames to target viewpoints achieves state-of-the-art long-term consistency and camera control.

  7. CustomX: Unified Character, Action, and Scene Customization in Video World Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    AniX generates controllable videos of a user-supplied character performing typed actions inside a user-supplied 3D scene by fine-tuning a pre-trained video generator on small locomotion datasets.

  8. PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention

    cs.CV 2025-11 conditional novelty 6.0 of 10

    PostCam generates new videos from a reference video along user-specified camera trajectories using a query-shared cross-attention that fuses pose data and rendered frames, improving control precision and detail preservation.

  9. Epipolar Geometry Improves Video Generation Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Ranking generated videos by their epipolar (Sampson) error and fine-tuning Wan2.1 with Flow-DPO cuts epipolar error 31% and raises human-rated 3D consistency from 54% to 72%.

  10. OmniNWM: Omniscient Driving Navigation World Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    OmniNWM jointly generates long panoramic multi-modal driving videos, controls them precisely via normalized Plücker ray-maps, and derives dense driving rewards from generated 3D occupancy.

  11. Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Hunyuan-GameCraft generates long, action-controlled game videos from a single image by unifying keyboard/mouse inputs into a continuous camera space and conditioning on mixed historical context.

  12. Video World Models with Long-term Spatial Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive video world model with a persistent static point-cloud spatial memory and sparse episodic keyframes improves revisit consistency over point-cloud-conditioned baselines.

  13. PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms

    cs.CV 2025-05 conditional novelty 6.0 of 10

    PanoWan adapts the Wan 2.1 text-to-video model to generate seamless 360-degree videos by remapping initial noise, rotating the latent grid during denoising, and padding the latent before VAE decoding, trained on a new...

  14. EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EPiC trains a 30M-parameter visibility-aware ControlNet on mask-based anchor videos from 5,000 in-the-wild videos and 500 steps, reaching SOTA camera accuracy on RealEstate10K and MiraData.

  15. EF-VI: Enhancing End-Frame Injection for Video Inbetweening

    cs.CV 2025-05 conditional novelty 6.0 of 10

    EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.

  16. PE-Field 4D: Video Generation Models as Canvas

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Warping reference tokens' positional encodings into the target view, with depth offsets and frame-level compression fixes, improves geometry-aware camera control in video diffusion transformers.

  17. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.

  18. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  19. From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

    cs.RO 2026-07 conditional novelty 4.0 of 10

    Physical intelligence needs an embodied brain that reasons over interventions and emits capability requests, grounded by a physical harness and shared experience contracts rather than direct actuator policies.

Pith tools