Pith. sign in

REVIEW 3 cited by

MatAnyone: Stable Video Matting with Consistent Memory Propagation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14677 v2 pith:ZJ5BYPUG submitted 2025-01-24 cs.CV

classification cs.CV
keywords mattingvideomemorymatanyonerobusttrainingconsistentdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To address this, we propose MatAnyone, a robust framework tailored for target-assigned video matting. Specifically, building on a memory-based paradigm, we introduce a consistent memory propagation module via region-adaptive memory fusion, which adaptively integrates memory from the previous frame. This ensures semantic stability in core regions while preserving fine-grained details along object boundaries. For robust training, we present a larger, high-quality, and diverse dataset for video matting. Additionally, we incorporate a novel training strategy that efficiently leverages large-scale segmentation data, boosting matting stability. With this new network design, dataset, and training strategy, MatAnyone delivers robust and accurate video matting results in diverse real-world scenarios, outperforming existing methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cinematic Compositing Using Character-Environment-Harmonized Video Generation Models

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    End-to-end video diffusion framework with tri-mask guidance and RGB-D denoising for joint modeling of character-to-environment physical interactions and environment-to-character lighting harmonization in cinematic com...

  2. BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The paper introduces a 3,254-video benchmark with pixel-level artifact masks for AI-generated video, and reports that fine-tuning on it improves artifact localization.

  3. IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

    cs.CV 2025-06 reject novelty 6.0 of 10

    A diffusion video model that jointly uses HDR lighting, relit frames, and 3D point tracks to relight videos from text prompts.

Pith tools