Pith. sign in

REVIEW 6 cited by

VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00522 v2 pith:M757D4R2 submitted 2024-03-01 cs.CV

classification cs.CV
keywords visionllamavisiontasksgenerationimagellamamanyprocess
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this question by unveiling a LLaMA-like vision transformer in plain and pyramid forms, termed VisionLLaMA, which is tailored for this purpose. VisionLLaMA is a unified and generic modelling framework for solving most vision tasks. We extensively evaluate its effectiveness using typical pre-training paradigms in a good portion of downstream tasks of image perception and especially image generation. In many cases, VisionLLaMA have exhibited substantial gains over the previous state-of-the-art vision transformers. We believe that VisionLLaMA can serve as a strong new baseline model for vision generation and understanding. Our code is released at https://github.com/Meituan-AutoML/VisionLLaMA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PixNerd: Pixel Neural Field Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PixNerd is a single-stage pixel-space diffusion transformer that uses predicted neural field weights to decode large patches, reaching 2.15 FID on ImageNet 256 without a VAE.

  2. Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    PyPE changes visual token position indices from raster order to concentric rings that flatten over layers, yielding small but broad accuracy gains across multiple VLM sizes and benchmarks.

  3. DiC: Rethinking Conv3x3 Designs in Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DiC, a U-shaped diffusion model using only 3x3 convolutions, sparse skips, and stage-wise conditioning, beats transformer-based DiT on ImageNet FID with higher throughput.

  4. Robust image classification with multi-modal large language models

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Multi-Shield rejects adversarial images by abstaining whenever a standard image classifier and a CLIP zero-shot classifier disagree, which raises robust accuracy noticeably under ordinary attacks and modestly under ad...

  5. GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Adding LLM-style GEGLU, RMSNorm, and rotary position embeddings to CoCa's vision encoder reduced contrastive loss, perplexity, and CoCa loss on one pretraining and three fine-tuning datasets, compared with an internal...

  6. UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A U-Net-style diffusion transformer (UDiT) is applied to quantum circuit synthesis, outperforming the U-Net-based GenQC on entanglement generation and unitary compilation in small-scale experiments.

Pith tools