Pith. sign in

REVIEW 5 cited by

InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.05778 v4 pith:QYEFJXBT submitted 2022-11-10 cs.CV

classification cs.CV
keywords internimagelarge-scalecnnsvitsmodelade20kcocodata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Compared to the great progress of large-scale vision transformers (ViTs) in recent years, large-scale models based on convolutional neural networks (CNNs) are still in an early state. This work presents a new large-scale CNN-based foundation model, termed InternImage, which can obtain the gain from increasing parameters and training data like ViTs. Different from the recent CNNs that focus on large dense kernels, InternImage takes deformable convolution as the core operator, so that our model not only has the large effective receptive field required for downstream tasks such as detection and segmentation, but also has the adaptive spatial aggregation conditioned by input and task information. As a result, the proposed InternImage reduces the strict inductive bias of traditional CNNs and makes it possible to learn stronger and more robust patterns with large-scale parameters from massive data like ViTs. The effectiveness of our model is proven on challenging benchmarks including ImageNet, COCO, and ADE20K. It is worth mentioning that InternImage-H achieved a new record 65.4 mAP on COCO test-dev and 62.9 mIoU on ADE20K, outperforming current leading CNNs and ViTs. The code will be released at https://github.com/OpenGVLab/InternImage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Attention Perturbations for Large Object Detection Transformers

    cs.CV 2025-08 conditional novelty 6.0 of 10

    AFOG, a learnable-attention white-box attack, collapses object detection mAP across twelve transformer detectors and three CNN detectors using small, visually imperceptible perturbations.

  2. Optimization of Collaborative Semantic Communication Network Performance with Channel and Content Preference Feedback

    cs.NI 2026-07 conditional novelty 5.0 of 10

    VDAC-DNC multi-agent RL jointly schedules sub-image channels and sparse CSI/content feedback, cutting semantic-weighted MSE by up to ~5–18% versus MAQAC and no-feedback baselines in simulation.

  3. LunarFM: A Shared Multimodal Representation of the Moon's Surface

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A self-supervised multimodal model fuses 18 channels from six lunar instruments into a shared 768-dimensional embedding per 0.5° chip, enabling mineral regression, similarity search, and geological-unit classification...

  4. GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy Prediction

    cs.CV 2025-07 conditional novelty 5.0 of 10

    GTAD combines an in-model latent denoising network with global temporal interaction to improve camera-based 3D semantic occupancy prediction, reporting 40.76 mIoU on Occ3D-nuScenes at 12 epochs.

  5. Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The ODOR dataset contributes 38,116 fine-grained object annotations over 4,712 artworks, benchmarked with five detector families, to stress-test object detection on dense, occluded, and off-centre objects in historica...

Pith tools