REVIEW 9 cited by
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating discrete tokens for text and continuous tokens for image. We find though there is an inherent trade-off between the image generation and understanding task, a carefully tuned training recipe enables them to improve each other. By selecting an appropriate loss balance weight, the unified model achieves results comparable to or exceeding those of single-task baselines on both tasks. Furthermore, we demonstrate that employing stronger pre-trained LLMs and random-order generation during training is important to achieve high-fidelity image generation within this unified framework. Built upon the Gemma model series, UniFluid exhibits competitive performance across both image generation and understanding, demonstrating strong transferability to various downstream tasks, including image editing for generation, as well as visual captioning and question answering for understanding.
Forward citations
Cited by 9 Pith papers
-
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
FeynmanBench is the first benchmark for evaluating multimodal LLMs on diagrammatic reasoning with Feynman diagrams, revealing systematic failures in enforcing physical constraints and global topology.
-
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
Adding depth- and segmentation-generation objectives to UMM post-training improved spatial understanding and reduced hallucinations on Harmon and OpenUni while preserving generation quality.
-
Reconstruction Alignment Improves Unified Multimodal Models
RECA, a self-supervised post-training objective that conditions unified multimodal models on their own visual understanding embeddings to reconstruct input images, improves text-to-image and editing benchmarks across ...
-
Generative Distribution Distillation
Knowledge distillation is reformulated as conditional diffusion over teacher feature tokens, with class-center contraction replacing the classification loss, yielding state-of-the-art ImageNet distillation numbers.
-
NeoBabel: A Multilingual Open Tower for Visual Generation
A 2B multilingual text-to-image model trained on 124M translated pairs matches or beats larger English-only baselines on English while scoring higher on the authors' multilingual benchmark extensions.
-
Fake it till You Make it: Reward Modeling as Discriminative Prediction
GAN-RM trains a CLIP-based discriminator to distinguish a few hundred preference proxy images from model outputs, then uses it for Best-of-N selection, SFT, and DPO.
-
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
A compact unified model that reuses a frozen VLM encoder and hybrid continuous/discrete tokens reaches competitive image understanding and generation with 15.6M training images and about $2,000 in compute.
-
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
A 1.5B unified autoregressive model with separate encoders for generation and understanding reports strong text-to-image and editing scores while running on commodity hardware.
-
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
A survey of instruction-based image editing plus a new 21-task benchmark, CDD-IIE, on which ten open models are scored by human experts.
Discussion (0). Sign in to comment.