Pith. sign in

REVIEW 10 cited by

EVA-02: A Visual Representation for Neon Genesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11331 v2 pith:MR5VZR5R submitted 2023-03-20 cs.CV cs.CL

classification cs.CVcs.CL
keywords eva-02parametersopenvisionaccessibleclipdataimagenet-1k
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We launch EVA-02, a next-generation Transformer-based visual representation pre-trained to reconstruct strong and robust language-aligned vision features via masked image modeling. With an updated plain Transformer architecture as well as extensive pre-training from an open & accessible giant CLIP vision encoder, EVA-02 demonstrates superior performance compared to prior state-of-the-art approaches across various representative vision tasks, while utilizing significantly fewer parameters and compute budgets. Notably, using exclusively publicly accessible training data, EVA-02 with only 304M parameters achieves a phenomenal 90.0 fine-tuning top-1 accuracy on ImageNet-1K val set. Additionally, our EVA-02-CLIP can reach up to 80.4 zero-shot top-1 on ImageNet-1K, outperforming the previous largest & best open-sourced CLIP with only ~1/6 parameters and ~1/6 image-text training data. We offer four EVA-02 variants in various model sizes, ranging from 6M to 304M parameters, all with impressive performance. To facilitate open access and open research, we release the complete suite of EVA-02 to the community at https://github.com/baaivision/EVA/tree/master/EVA-02.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MentalThink: Shaping Thoughts in Mental SVG World

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.

  2. GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A gated, training-free local-intrinsic-dimension profile from a frozen ViT repairs face-forgery detectors on unseen GAN and diffusion axes, lifting generation-family AUC by +0.084.

  3. Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A continuously refreshed, incentive-driven deepfake detector beats static detectors on in-the-wild benchmarks and improves on post-export AI-generated media.

  4. ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    EVT improves the RMT backbone by using Euclidean-distance attention decay and 1D token grouping, achieving 86.6% top-1 on ImageNet-1K at 384×384 resolution.

  5. VoCap: Video Object Captioning and Segmentation from Any Prompt

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VoCap jointly performs promptable video object segmentation and object captioning, and introduces a 50k-video pseudo-caption dataset that improves both tasks.

  6. A Generalized Learning Framework for Self-Supervised Contrastive Learning

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A single framework unifies BYOL, Barlow Twins, and SwAV, plus a plug-in calibration method, ADC, that improves learned representations by preserving input-space distances.

  7. Multi-Granularity Feature Calibration via VFM for Domain Generalized Semantic Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    MGFC calibrates VFM features at coarse, medium, and fine granularity to achieve state-of-the-art domain generalized semantic segmentation on GTA5-to-real and Cityscapes-to-ACDC benchmarks.

  8. EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

    cs.CV 2025-06 conditional novelty 5.0 of 10

    EVA02-AT combines full-dimension spatial and temporal rotary position embeddings with a symmetric multi-similarity loss to improve egocentric video-text retrieval.

  9. Generalized Trajectory Scoring for End-to-end Multimodal Planning

    cs.RO 2025-06 conditional novelty 5.0 of 10

    GTRS combines super-dense vocabulary training, dropout, sensor augmentation, and diffusion proposals to reach 49.4 EPDMS on the Navhard benchmark, approaching the privileged PDM-Closed method.

  10. Twistronics and moir\'e superlattice physics in 2D transition metal dichalcogenides

    cond-mat.mes-hall 2025-08 unverdicted

    The provided full text does not match the abstract, so the review's content cannot be assessed.

Pith tools