Pith. sign in

REVIEW 1 cited by

Object Isolated Attention for Consistent Story Visualization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.23353 v1 pith:VR5NTDTD submitted 2025-03-30 cs.CV cs.AI

classification cs.CVcs.AI
keywords attentioncharacterconsistencyisolatedcrossfeaturesmechanismmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character consistency while creating natural and contextually fitting scenes--an area where many existing methods struggle. In this paper, we propose an enhanced Transformer module that uses separate self attention and cross attention mechanisms, leveraging prior knowledge from pre-trained diffusion models to ensure logical scene creation. The isolated self attention mechanism improves character consistency by refining attention maps to reduce focus on irrelevant areas and highlight key features of the same character. Meanwhile, the isolated cross attention mechanism independently processes each character's features, avoiding feature fusion and further strengthening consistency. Notably, our method is training-free, allowing the continuous generation of new characters and storylines without re-tuning. Both qualitative and quantitative evaluations show that our approach outperforms current methods, demonstrating its effectiveness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unpaired Deblurring via Decoupled Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A diffusion model that decouples structural features from blur patterns using unpaired target-domain images can deblur photos in unseen domains without paired training data.

Pith tools