REVIEW 3 cited by
MatAnyone: Stable Video Matting with Consistent Memory Propagation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To address this, we propose MatAnyone, a robust framework tailored for target-assigned video matting. Specifically, building on a memory-based paradigm, we introduce a consistent memory propagation module via region-adaptive memory fusion, which adaptively integrates memory from the previous frame. This ensures semantic stability in core regions while preserving fine-grained details along object boundaries. For robust training, we present a larger, high-quality, and diverse dataset for video matting. Additionally, we incorporate a novel training strategy that efficiently leverages large-scale segmentation data, boosting matting stability. With this new network design, dataset, and training strategy, MatAnyone delivers robust and accurate video matting results in diverse real-world scenarios, outperforming existing methods.
Forward citations
Cited by 3 Pith papers
-
Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models
End-to-end video diffusion framework with tri-mask guidance and RGB-D denoising for joint modeling of character-to-environment physical interactions and environment-to-character lighting harmonization in cinematic com...
-
BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos
The paper introduces a 3,254-video benchmark with pixel-level artifact masks for AI-generated video, and reports that fine-tuning on it improves artifact localization.
-
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
A diffusion video model that jointly uses HDR lighting, relit frames, and 3D point tracks to relight videos from text prompts.
Discussion (0). Continue with ORCID to comment.