Pith. sign in

REVIEW 21 cited by

P+: Extended Textual Conditioning in Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.09522 v3 pith:3L3SIAOR submitted 2023-03-16 cs.CV cs.CLcs.GRcs.LG

classification cs.CVcs.CLcs.GRcs.LG
keywords spaceextendedtextualtext-to-imageinversionmodelsconditioningintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce an Extended Textual Conditioning space in text-to-image models, referred to as $P+$. This space consists of multiple textual conditions, derived from per-layer prompts, each corresponding to a layer of the denoising U-net of the diffusion model. We show that the extended space provides greater disentangling and control over image synthesis. We further introduce Extended Textual Inversion (XTI), where the images are inverted into $P+$, and represented by per-layer tokens. We show that XTI is more expressive and precise, and converges faster than the original Textual Inversion (TI) space. The extended inversion method does not involve any noticeable trade-off between reconstruction and editability and induces more regular inversions. We conduct a series of extensive experiments to analyze and understand the properties of the new space, and to showcase the effectiveness of our method for personalizing text-to-image models. Furthermore, we utilize the unique properties of this space to achieve previously unattainable results in object-style mixing using text-to-image models. Project page: https://prompt-plus.github.io

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers

    cs.CV 2025-05 conditional novelty 7.0 of 10

    LoRAShop localizes each LoRA's effect to attention-derived spatial masks inside a Flux transformer, enabling training-free multi-concept image generation and editing.

  2. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.

  3. PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Two-stage text-guided mesh deformation (Laplacian CLIP scaling + attention-shared SDS Jacobian sculpting) better preserves source pose while aligning to text than TextDeformer or MeshUp.

  4. LILAC: Layer-Wise Independent LoRAs and Cascaded Conditioning for Multi-Concept Customization of Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Independently trained LoRAs composed as sequential layers with frozen conditioning preserve multi-subject identity better than weight-space fusion, reaching 0.861 ArcFace detection rate.

  5. Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    Implicit generative choices in diffusion models for ambiguous prompts are localized principally in self-attention layers, enabling a targeted ICM steering method that outperforms prior debiasing approaches.

  6. GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization

    cs.CV 2026-01 conditional novelty 6.0 of 10

    GimmBO uses preference-based Bayesian optimization with a sparse, sum-bounded search space to help users interactively discover adapter merges in 20-30 dimensional model-merging spaces.

  7. Alterbute: Editing Intrinsic Attributes of Objects in Images

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Alterbute performs identity-preserving editing of an object's intrinsic attributes (color, texture, material, shape) using Visual-Named-Entity-based identity supervision and a relaxed training objective.

  8. Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.

  9. Per-Query Visual Concept Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A prompt- and seed-specific, attention-based loss step improves both identity preservation and prompt adherence for six personalization methods across SD, SDXL, and FLUX backbones.

  10. TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    TARA adds token-focused masking and a token alignment loss to LoRA adapters, allowing several independently trained personalized adapters to be composed with less identity loss and feature leakage.

  11. SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    SCFlow learns a reversible style-content merge and then lets the same mapping perform separation without explicit disentanglement training.

  12. APT: Adaptive Personalized Training for Diffusion Models with Limited Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    APT detects overfitting during diffusion fine-tuning and uses adaptive augmentation, loss weighting, feature-statistics regularization, and attention alignment to preserve prior knowledge while learning new concepts.

  13. Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Difference Inversion learns text tokens encoding the A-to-A' edit and applies those tokens to B to synthesize B' with Stable Diffusion, without model-specific tuning.

  14. Noise Consistency Regularization for Improved Subject-Driven Image Synthesis

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Adding consistency-to-pretrained and multiplicative-noise consistency losses to fine-tuning improves subject identity and background diversity over DreamBooth on a 30-subject benchmark.

  15. MARBLE: Material Recomposition and Blending in CLIP-Space

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MARBLE performs material blending and parametric material-attribute control by manipulating CLIP image embeddings and injecting them into a specific U-Net block of a pre-trained diffusion model.

  16. AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    AlginGen improves zero-shot personalized image generation by training a learnable token and a selective attention mask that align textual and visual priors, achieving the best balance of concept preservation and promp...

  17. CDST: Color Disentangled Style Transfer for Universal Style Reference Customization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CDST disentangles color from style via greyscale style input and a color histogram stream, enabling zero-shot style transfer with separate color control and a new characteristics-preserved mode.

  18. Bridging Rendering and Generative Modeling with Monte Carlo Transport Scheduling

    cs.CV 2026-02 reject novelty 5.0 of 10

    A common variance-time SDE aligns Monte Carlo rendering noise with diffusion-model denoising, enabling low-spp render refinement and stage-ordered material control.

  19. From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Wardrobe Polyptych LoRA lets a single diffusion model compose a person's face and clothing from multiple reference photos into new full-body images, generalizing to unseen identities without inference-time fine-tuning.

  20. Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Re-centering and re-scaling the consistency guidance component parallel to the text direction improves prompt adherence with only a small drop in identity preservation.

  21. Training Free Stylized Abstraction

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A training-free framework coupling VLLM-based identity distillation with cross-domain rectified flow inversion generates identity-preserving stylized abstractions from a single reference image, evaluated by a new GPT-...

Pith tools