Pith. sign in

REVIEW 3 cited by

Diffutoon: High-Resolution Editable Toon Shading via Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.16224 v1 pith:UR7E2XQ5 submitted 2024-01-29 cs.CV

classification cs.CV
keywords diffutoonshadingtoondiffusionmodelsstylizationvideosanime
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Toon shading is a type of non-photorealistic rendering task of animation. Its primary purpose is to render objects with a flat and stylized appearance. As diffusion models have ascended to the forefront of image synthesis methodologies, this paper delves into an innovative form of toon shading based on diffusion models, aiming to directly render photorealistic videos into anime styles. In video stylization, extant methods encounter persistent challenges, notably in maintaining consistency and achieving high visual quality. In this paper, we model the toon shading problem as four subproblems: stylization, consistency enhancement, structure guidance, and colorization. To address the challenges in video stylization, we propose an effective toon shading approach called \textit{Diffutoon}. Diffutoon is capable of rendering remarkably detailed, high-resolution, and extended-duration videos in anime style. It can also edit the content according to prompts via an additional branch. The efficacy of Diffutoon is evaluated through quantitive metrics and human evaluation. Notably, Diffutoon surpasses both open-source and closed-source baseline approaches in our experiments. Our work is accompanied by the release of both the source code and example videos on Github (Project page: https://ecnu-cilab.github.io/DiffutoonProjectPage/).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A 3D-VQGAN with spatial-temporal codebooks and code-lookup transformers restores compressed face videos and removes flicker in about 3 seconds per 24-frame clip.

  2. EndoControlMag: Robust Endoscopic Vascular Motion Magnification with Periodic Reference Resetting and Hierarchical Tissue-aware Dual-Mask Control

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A training-free Lagrangian motion magnification framework with periodic reference resetting and tissue-aware dual-mask control improves vascular pulsation visibility in endoscopic surgery videos.

  3. Single Trajectory Distillation for Accelerating Image and Video Style Transfer

    cs.CV 2024-12 conditional novelty 4.0 of 10

    A consistency-distillation method trains from a fixed partial-noise start and uses a replay bank plus DINO-v2 adversarial loss to improve few-step image and video stylization.

Pith tools