Pith. sign in

REVIEW 4 cited by

Sakuga-42M Dataset: Scaling Up Cartoon Research

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.07425 v1 pith:36PYDPEV submitted 2024-05-13 cs.CV

classification cs.CV
keywords cartoondatasetresearchmodelssakuga-42mscalingvideoanimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural video by scaling large models with extensive datasets, they are not as effective for cartoons. Through our empirical experiments, we argue that this ineffectiveness stems from a notable bias in hand-drawn cartoons that diverges from the distribution of natural videos. Can we harness the success of the scaling paradigm to benefit cartoon research? Unfortunately, until now, there has not been a sizable cartoon dataset available for exploration. In this research, we propose the Sakuga-42M Dataset, the first large-scale cartoon animation dataset. Sakuga-42M comprises 42 million keyframes covering various artistic styles, regions, and years, with comprehensive semantic annotations including video-text description pairs, anime tags, content taxonomies, etc. We pioneer the benefits of such a large-scale cartoon dataset on comprehension and generation tasks by finetuning contemporary foundation models like Video CLIP, Video Mamba, and SVD, achieving outstanding performance on cartoon-related tasks. Our motivation is to introduce large-scaling to cartoon research and foster generalization and robustness in future cartoon applications. Dataset, Code, and Pretrained Models will be publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free inference framework combining target-aware reference expansion, soft top-k palette voting, and cycle-gated temporal fusion improves segment-matching colourisation on animation videos.

  2. MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MagicAnime is a 400k-clip multimodal cartoon dataset with hierarchical annotations and benchmarks for image-to-video, pose-driven, face reenactment, and audio-driven animation generation.

  3. LongAnimation: Long Animation Generation with Dynamic Global-Local Memory

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LongAnimation uses a dynamic global-local memory, built from a long-video-understanding model's KV cache, to colorize animation sequences of about 500 frames with stable color consistency.

  4. SketchColour: Channel Concat Guided DiT-based Sketch-to-Colour Pipeline for 2D Animation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SketchColour colors animation sketches from a single colored first frame by replacing the U-Net with a Diffusion Transformer, using channel-concat conditioning and LoRA fine-tuning.

Pith tools