Pith. sign in

REVIEW 8 cited by

StyleShot: A Snapshot on Any Style

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01414 v2 pith:5Z3RORPZ submitted 2024-07-01 cs.CV

classification cs.CV
keywords stylestyleshotencoderstylesrepresentationstyle-awarestylegallerytest-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we show that, a good style representation is crucial and sufficient for generalized style transfer without test-time tuning. We achieve this through constructing a style-aware encoder and a well-organized style dataset called StyleGallery. With dedicated design for style learning, this style-aware encoder is trained to extract expressive style representation with decoupling training strategy, and StyleGallery enables the generalization ability. We further employ a content-fusion encoder to enhance image-driven style transfer. We highlight that, our approach, named StyleShot, is simple yet effective in mimicking various desired styles, i.e., 3D, flat, abstract or even fine-grained styles, without test-time tuning. Rigorous experiments validate that, StyleShot achieves superior performance across a wide range of styles compared to existing state-of-the-art methods. The project page is available at: https://styleshot.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.

  2. USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.

  3. Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.

  4. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  5. MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MotionShot transfers motion from a reference video to an unseen target object in text-to-video generation by combining semantic and morphological alignment in a training-free pipeline.

  6. One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A single adversarial image can make a unified vision-language model misclassify the same object across captioning, detection, region classification, and localization, and the new CrossVLAD benchmark and CRAFT attack m...

  7. OmniStyle: Filtering High Quality Style Transfer Data at Scale

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new million-triplet dataset and a diffusion transformer model that performs text-guided and image-guided style transfer, with a filtering pipeline used to curate high-quality training examples.

  8. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

Pith tools