Pith. sign in

REVIEW 3 cited by

Importance-Based Token Merging for Efficient Image and Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16720 v2 pith:L77RHSW2 submitted 2024-11-23 cs.CV

classification cs.CV
keywords mergingtokengenerationtokensacrossapproachdetailsdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-information tokens during merging - those essential for semantic fidelity and structural details - significantly improves sample quality, producing finer details and more coherent, realistic generations. Despite being simple and intuitive, this approach remains underexplored. To do so, we propose an importance-based token merging method that prioritizes the most critical tokens in computational resource allocation, leveraging readily available importance scores, such as those from classifier-free guidance in diffusion models. Experiments show that our approach significantly outperforms baseline methods across multiple applications, including text-to-image synthesis, multi-view image generation, and video generation with various model architectures such as Stable Diffusion, Zero123++, AnimateDiff, or PixArt-$\alpha$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Local Representative Token Guided Merging for Text-to-Image Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ReToM merges tokens around a similarity-selected representative token in adaptive local windows, improving Stable Diffusion FID from 37.02 to 34.89 at comparable inference speed.

  2. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  3. RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Training-free sparse attention that classifies each head as spatial, temporal, or textural and applies a matched mask or token reduction, giving about 1.9x attention speedup with small VBench losses.

Pith tools