REVIEW 6 cited by
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA.
Forward citations
Cited by 6 Pith papers
-
PixNerd: Pixel Neural Field Diffusion
PixNerd is a single-stage pixel-space diffusion transformer that uses predicted neural field weights to decode large patches, reaching 2.15 FID on ImageNet 256 without a VAE.
-
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
PyPE changes visual token position indices from raster order to concentric rings that flatten over layers, yielding small but broad accuracy gains across multiple VLM sizes and benchmarks.
-
DiC: Rethinking Conv3x3 Designs in Diffusion Models
DiC, a U-shaped diffusion model using only 3x3 convolutions, sparse skips, and stage-wise conditioning, beats transformer-based DiT on ImageNet FID with higher throughput.
-
Robust image classification with multi-modal large language models
Multi-Shield rejects adversarial images by abstaining whenever a standard image classifier and a CLIP zero-shot classifier disagree, which raises robust accuracy noticeably under ordinary attacks and modestly under ad...
-
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internal...
-
UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis
A U-Net-style diffusion transformer (UDiT) is applied to quantum circuit synthesis, outperforming the U-Net-based GenQC on entanglement generation and unitary compilation in small-scale experiments.
Discussion (0). Continue with ORCID to comment.