Pith. sign in

REVIEW 10 cited by

Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02308 v3 pith:FAXCLRDQ submitted 2024-03-04 cs.CV

classification cs.CV
keywords processinghigh-resolutionmodeltasksvisionvision-rwkvvrwkvcomplexity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model adapted from the RWKV model used in the NLP field with necessary modifications for vision tasks. Similar to the Vision Transformer (ViT), our model is designed to efficiently handle sparse inputs and demonstrate robust global processing capabilities, while also scaling up effectively, accommodating both large-scale parameters and extensive datasets. Its distinctive advantage lies in its reduced spatial aggregation complexity, which renders it exceptionally adept at processing high-resolution images seamlessly, eliminating the necessity for windowing operations. Our evaluations demonstrate that VRWKV surpasses ViT's performance in image classification and has significantly faster speeds and lower memory usage processing high-resolution inputs. In dense prediction tasks, it outperforms window-based models, maintaining comparable speeds. These results highlight VRWKV's potential as a more efficient alternative for visual perception tasks. Code is released at https://github.com/OpenGVLab/Vision-RWKV.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. pLSTM: parallelizable Linear Source Transition Mark networks

    cs.LG 2025-06 conditional novelty 7.0 of 10

    pLSTM extends linear recurrent networks to general directed acyclic graphs with a parallelizable scheme and two stabilization modes for long-range propagation.

  2. LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Dual-stream HDR video can be reconstructed without fragile cross-exposure warping by backbone-guided collaborative attention plus sequence-level residual refinement.

  3. AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition

    cs.SD 2025-09 conditional novelty 6.0 of 10

    AudioRWKV, an RWKV7-based audio backbone with 2D convolution and bidirectional WKV, beats AST and AuM baselines on five benchmarks at linear complexity.

  4. PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PointDGRWKV applies RWKV-like attention to domain-generalized point cloud classification, adding a geometric token shift and key-distribution alignment, and reports state-of-the-art accuracy on PointDA-10 and PointDG-3to1.

  5. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  6. FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FEAT combines spatial, temporal, and channel attention in a diffusion transformer and reports improved medical video generation metrics with lower parameter counts than Endora.

  7. URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restoration

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A multi-state RWKV-based model, with luminance-adaptive normalization and state-aware selective fusion, reports improved PSNR and SSIM across six low-light benchmarks.

  8. EfficientIML: Efficient High-Resolution Image Manipulation Localization

    cs.CV 2025-09 conditional novelty 5.0 of 10

    EfficientIML combines a lightweight RWKV-based backbone with multi-scale supervision to accurately localize high-resolution diffusion-based image forgeries, while introducing the SIF dataset for training and evaluation.

  9. U-RWKV: Lightweight medical image segmentation with direction-adaptive RWKV

    eess.IV 2025-07 conditional novelty 5.0 of 10

    U-RWKV is a lightweight U-shaped medical image segmenter that combines multi-directional RWKV scanning with stage-adaptive channel recalibration, reporting competitive Dice scores with about three million parameters.

  10. Med-URWKV{\dag}: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation

    eess.IV 2025-06 conditional novelty 5.0 of 10

    Pretrained pure VRWKV encoders paired with pure VRWKV decoders match or beat CNN, ViT, and Mamba baselines, with a small model plus FAWA and MSCF modules reaching 88% average Dice.

Pith tools