REVIEW 10 cited by
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processing and long-context analysis. This paper introduces Vision-RWKV (VRWKV), a model adapted from the RWKV model used in the NLP field with necessary modifications for vision tasks. Similar to the Vision Transformer (ViT), our model is designed to efficiently handle sparse inputs and demonstrate robust global processing capabilities, while also scaling up effectively, accommodating both large-scale parameters and extensive datasets. Its distinctive advantage lies in its reduced spatial aggregation complexity, which renders it exceptionally adept at processing high-resolution images seamlessly, eliminating the necessity for windowing operations. Our evaluations demonstrate that VRWKV surpasses ViT's performance in image classification and has significantly faster speeds and lower memory usage processing high-resolution inputs. In dense prediction tasks, it outperforms window-based models, maintaining comparable speeds. These results highlight VRWKV's potential as a more efficient alternative for visual perception tasks. Code is released at https://github.com/OpenGVLab/Vision-RWKV.
Forward citations
Cited by 10 Pith papers
-
pLSTM: parallelizable Linear Source Transition Mark networks
pLSTM extends linear recurrent networks to general directed acyclic graphs with a parallelizable scheme and two stabilization modes for long-range propagation.
-
LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction
Dual-stream HDR video can be reconstructed without fragile cross-exposure warping by backbone-guided collaborative attention plus sequence-level residual refinement.
-
AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
AudioRWKV, an RWKV7-based audio backbone with 2D convolution and bidirectional WKV, beats AST and AuM baselines on five benchmarks at linear complexity.
-
PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification
PointDGRWKV applies RWKV-like attention to domain-generalized point cloud classification, adding a geometric token shift and key-distribution alignment, and reports state-of-the-art accuracy on PointDA-10 and PointDG-3to1.
-
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.
-
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
FEAT combines spatial, temporal, and channel attention in a diffusion transformer and reports improved medical video generation metrics with lower parameter counts than Endora.
-
URWKV: Unified RWKV Model with Multi-state Perspective for Low-light Image Restoration
A multi-state RWKV-based model, with luminance-adaptive normalization and state-aware selective fusion, reports improved PSNR and SSIM across six low-light benchmarks.
-
EfficientIML: Efficient High-Resolution Image Manipulation Localization
EfficientIML combines a lightweight RWKV-based backbone with multi-scale supervision to accurately localize high-resolution diffusion-based image forgeries, while introducing the SIF dataset for training and evaluation.
-
U-RWKV: Lightweight medical image segmentation with direction-adaptive RWKV
U-RWKV is a lightweight U-shaped medical image segmenter that combines multi-directional RWKV scanning with stage-adaptive channel recalibration, reporting competitive Dice scores with about three million parameters.
-
Med-URWKV{\dag}: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation
Pretrained pure VRWKV encoders paired with pure VRWKV decoders match or beat CNN, ViT, and Mamba baselines, with a small model plus FAWA and MSCF modules reaching 88% average Dice.
Discussion (0). Continue with ORCID to comment.