Pith. sign in

REVIEW 3 cited by

Splatter a Video: Video Gaussian Representation for Versatile Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13870 v2 pith:7ACGDLZL submitted 2024-06-19 cs.CV

classification cs.CV
keywords videorepresentationdepthexplicitgaussiantasksappearanceediting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video representation is a long-standing problem that is crucial for various down-stream tasks, such as tracking,depth prediction,segmentation,view synthesis,and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D representations that are ill-suited for manipulation tasks. To address these challenges, we introduce a novel explicit 3D representation-video Gaussian representation -- that embeds a video into 3D Gaussians. Our proposed representation models video appearance in a 3D canonical space using explicit Gaussians as proxies and associates each Gaussian with 3D motions for video motion. This approach offers a more intrinsic and explicit representation than layered atlas or volumetric pixel matrices. To obtain such a representation, we distill 2D priors, such as optical flow and depth, from foundation models to regularize learning in this ill-posed setting. Extensive applications demonstrate the versatility of our new video representation. It has been proven effective in numerous video processing tasks, including tracking, consistent video depth and feature refinement, motion and appearance editing, and stereoscopic video generation. Project page: https://sunyangtian.github.io/spatter_a_video_web/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GSVR: 2D Gaussian-based Video Representation for 800+ FPS with Hybrid Deformation Field

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2D Gaussian video representation with a tri-plane plus polynomial deformation field decodes at 800+ FPS on Bunny and trains in about 2 seconds per frame.

  2. EX-4D: EXtreme Viewpoint 4D Video Synthesis via Depth Watertight Mesh

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EX-4D uses a depth watertight mesh and simulated occlusion masks to condition a video diffusion model for extreme-viewpoint 4D video synthesis from monocular input.

  3. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

Pith tools