Pith. sign in

REVIEW 12 cited by

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06135 v1 pith:WMYI6KBX submitted 2024-07-08 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords generationmultimodalanolelargemodelsnativeautoregressiveimage-text
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous open-source large multimodal models (LMMs) have faced several limitations: (1) they often lack native integration, requiring adapters to align visual representations with pre-trained large language models (LLMs); (2) many are restricted to single-modal generation; (3) while some support multimodal generation, they rely on separate diffusion models for visual modeling and generation. To mitigate these limitations, we present Anole, an open, autoregressive, native large multimodal model for interleaved image-text generation. We build Anole from Meta AI's Chameleon, adopting an innovative fine-tuning strategy that is both data-efficient and parameter-efficient. Anole demonstrates high-quality, coherent multimodal generation capabilities. We have open-sourced our model, training framework, and instruction tuning data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VIG-RL: Learning to Search and Insert for Verified Image Grounding

    cs.IR 2026-07 conditional novelty 6.0 of 10

    An RL-trained ReAct agent learns joint search–selection–insertion of retrieved authentic images, setting SOTA on MRAMG-Bench Verified Image Grounding.

  2. Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    p-less cluster decoding, which truncates and samples over K-means clusters of visual tokens rather than individual tokens, yields higher per-prompt sample diversity than default or dynamic-temperature baselines on mos...

  3. DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

    cs.LG 2026-07 accept novelty 6.0 of 10

    Schema-guided interleaved state-transition pretraining with selective attention and reweighted loss improves hierarchical visual dynamics modeling for narrative generation and world simulation.

  4. Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Supervising text–image handoffs with Reflective SFT and Flow-GRPO (MoTiF) reduces modal isolation and raises accuracy on four visual puzzle benchmarks versus end-task-only training.

  5. SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

    cs.CV 2026-03 accept novelty 6.0 of 10

    SJD-PAC combines proactive multi-path drafting and adaptive continuation to raise average acceptance length in Speculative Jacobi Decoding, delivering 3.8 imes lossless wall-clock speedup on Lumina-mGPT and Emu3.

  6. UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    UniCode² builds a 500K-entry codebook from clustered SigLIP embeddings and uses a cascaded frozen-plus-trainable codebook to unify multimodal understanding and generation with stable training and high token utilization.

  7. Can VLMs Predict Future States? Bootstrapping World Models from Inverse Dynamics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using an inverse dynamics model to label and verify frame transitions lets vision-language models learn forward dynamics and outperform specialized image editors on Aurora-Bench.

  8. Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    On the ACON benchmark, any-to-any models do not consistently beat specialist model pairs on cyclic consistency, but show weak latent-space consistency in equivariance tests.

  9. ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Progressive multi-step latent visual thoughts, endogenously distilled from a model's own encoder on synthetic trajectories and regularized by distance-weighted diversity, improve MLLM visual reasoning accuracy and efficiency.

  10. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  11. Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Selecting the intermediate layer where image-conditioned and text-only predictions diverge most, and adding that layer's contrastive visual signal back to the final logits, reduces object hallucinations in four large ...

  12. StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation

    cs.CV 2025-05 reject novelty 4.0 of 10

    StyleAR enables autoregressive image generation models to do style-aligned text-to-image generation using only binary text-image data, via self-reconstruction training and style-enhanced tokens.

Pith tools