Pith. sign in

REVIEW 31 cited by

In-Context LoRA for Diffusion Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23775 v3 pith:RCKAOG7G submitted 2024-10-31 cs.CV cs.GR

classification cs.CVcs.GR
keywords in-contexttuningditsgenerationimagesdataloramodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent research arXiv:2410.15027 has explored the use of diffusion transformers (DiTs) for task-agnostic image generation by simply concatenating attention tokens across images. However, despite substantial computational resources, the fidelity of the generated images remains suboptimal. In this study, we reevaluate and streamline this framework by hypothesizing that text-to-image DiTs inherently possess in-context generation capabilities, requiring only minimal tuning to activate them. Through diverse task experiments, we qualitatively demonstrate that existing text-to-image DiTs can effectively perform in-context generation without any tuning. Building on this insight, we propose a remarkably simple pipeline to leverage the in-context abilities of DiTs: (1) concatenate images instead of tokens, (2) perform joint captioning of multiple images, and (3) apply task-specific LoRA tuning using small datasets (e.g., 20~100 samples) instead of full-parameter tuning with large datasets. We name our models In-Context LoRA (IC-LoRA). This approach requires no modifications to the original DiT models, only changes to the training data. Remarkably, our pipeline generates high-fidelity image sets that better adhere to prompts. While task-specific in terms of tuning data, our framework remains task-agnostic in architecture and pipeline, offering a powerful tool for the community and providing valuable insights for further research on product-level task-agnostic generation systems. We release our code, data, and models at https://github.com/ali-vilab/In-Context-LoRA

Discussion (0). Sign in to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation

    cs.CV 2026-07 accept novelty 7.0 of 10

    CtrlVTON recasts virtual try-on as mask-conditioned editing and introduces VIP-SAM for instance-level garment segmentation, beating proprietary editors on layout fidelity while matching garment quality.

  2. MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    MAVIN proposes boundary-aware attention, ID-aware propagation, a multi-agent scripting pipeline, and the MAVINSet dataset as the first framework for multi-shot audio-visual generation with narrative control, claiming ...

  3. WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing

    cs.CV 2026-03 conditional novelty 7.0 of 10

    WeEdit trains a glyph-guided, RL-optimized image editor on a 330K-pair synthetic multilingual dataset and reports open-source SOTA on its own bilingual and multilingual text-editing benchmarks.

  4. Story2Board: A Training-Free Approach for Expressive Storyboard Generation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    Story2Board uses reciprocal attention value mixing and latent panel anchoring to generate consistent yet visually diverse storyboards from text without any training.

  5. Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

    cs.CV 2025-06 conditional novelty 7.0 of 10

    IT-Blender blends a real image and a text prompt by injecting clean reference features into a trained attention module, improving disentangled concept blending in SD and FLUX.

  6. DynEval: Holistic Evaluations of T2I Generative Models in the Wild

    cs.CV 2026-07 conditional novelty 6.5 of 10

    DynEval distills a 235B teacher VLM into 2B/4B evaluators via 250K synthetic instruction triplets, yielding higher human correlation than existing T2I metrics while enabling open-set dynamic QA and scene-graph quality checks.

  7. Global Pose Control for Generative View Synthesis in Normalized Object Coordinate Space

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A diffusion image-editing model conditioned on Plücker ray-map tokens and text-defined NOCS fronts generates high-fidelity novel views with absolute global pose control from unposed inputs.

  8. VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An MLLM extracts transferable physical cues from a reference video and conditions a pretrained I2V generator so new scenes follow that physics without exhaustive prompts.

  9. LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A three-stage pipeline decomposes indoor scene lighting into individually controllable OLAT sources, harmonizes the decomposition across views, and encodes it in 3D Gaussian splatting for real-time per-light editing.

  10. TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning

    cs.CV 2025-12 conditional novelty 6.0 of 10

    TinyHistory compresses long video history into a ~5k-token context via a two-stage learning scheme, achieving consistency on par with heavier baselines at lower memory cost.

  11. UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A reinforcement-learning reward based on bipartite face matching improves multi-identity consistency and reduces identity confusion in image customization models.

  12. FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FocusDPO adds dynamic spatial weighting to preference-based fine-tuning, improving subject fidelity and reducing attribute leakage in multi-subject personalized image generation.

  13. USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.

  14. ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    ToonComposer generates cartoon videos from a colored reference frame and sparse keyframe sketches, merging inbetweening and colorization in one diffusion model.

  15. AnimeColor: Reference-based Animation Colorization with Diffusion Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnimeColor colorizes animation sketch sequences from a reference image using a diffusion transformer with high-level and low-level color guidance.

  16. FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.

  17. ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A panorama representation and attention scheme that lets a pretrained perspective video diffusion model generate spatially consistent 360-degree videos from an input perspective clip.

  18. RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RecipeGen is a new benchmark with 26,453 recipes, 196,724 step-aligned images, and 4,491 cooking videos, plus three domain-specific evaluation metrics for recipe generation.

  19. IA-T2I: Internet-Augmented Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    IA-T2I uses active retrieval, hierarchical image selection, and self-reflection to supply internet reference images to T2I models, improving generation accuracy on uncertain-knowledge prompts.

  20. ShoulderShot: Generating Over-the-Shoulder Dialogue Videos

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ShoulderShot generates over-the-shoulder dialogue videos by pairing two linked camera shots and looping them, so characters stay consistent through long multi-turn conversations.

  21. ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

    cs.CV 2025-08 reject novelty 5.0 of 10

    ICM-Fusion uses a conditional VAE plus task-vector guidance to fuse multiple LoRA adapters into one model, reporting marginal average gains on vision and language benchmarks and larger gains in a few-shot long-tail setup.

  22. Steering Guidance for Personalized Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Weight-interpolated null-text weak model in classifier-free guidance improves subject fidelity with minimal text-fidelity loss in personalized text-to-image diffusion.

  23. Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits

    physics.optics 2025-08 reject novelty 5.0 of 10

    The abstract reports a low-loss ScAlN/Si3N4 hybrid waveguide, but the full text is an unrelated diffusion-model paper, leaving the photonics claim without supporting evidence.

  24. Captain Cinema: Towards Short Movie Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A two-stage text-to-movie system that plans keyframes for the story and then synthesizes video between them, using a compressed memory bank to keep long narratives consistent.

  25. From Wardrobe to Canvas: Wardrobe Polyptych LoRA for Part-level Controllable Human Image Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Wardrobe Polyptych LoRA lets a single diffusion model compose a person's face and clothing from multiple reference photos into new full-body images, generalizing to unseen identities without inference-time fine-tuning.

  26. PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.

  27. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  28. Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.

  29. In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    In-Context Brush performs zero-shot customized subject insertion by amplifying prompt and reference attention and reweighting attention heads in a pre-trained Flux-Fill diffusion transformer.

  30. Hunyuan-Game: Industrial-grade Intelligent Game Creation Model

    cs.CV 2025-05 reject novelty 4.0 of 10

    Tencent's Hunyuan-Game applies diffusion transformers to game asset creation across nine image and video generation tasks, with self-reported gains that are partly contradicted by its own evaluation table.

  31. Fuel Consumption in Platoons: A Literature Review

    eess.SY 2025-08 unverdicted

    A literature review compiling factors that affect fuel consumption in vehicle platoons, including drag reduction, coordination, and instability.

Pith tools