REVIEW 3 major objections 7 minor 62 references
MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection
T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read MamKPD, a Mamba pose model, reaches 77.3% AP at 1492 FPS on COCO.
desk verdict A genuinely first Mamba-based 2D pose estimator with real parameter efficiency, but the headline speed claim rests on unmatched comparisons and needs fairer benchmarks and code before it can be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the contextual modeling module (CMM) used inside every stage. CMM is a lightweight two-part block: a depth-wise convolution along the patch dimension, followed by a patch-embedding layer and normalization, models dependencies between patches and expands the receptive field; a linear layer plus a 1x1 depth-wise convolution then distills pose cues within each patch. This is interleaved with the SS2D block, which scans the feature map in four directions to aggregate global patch information. The combination is what lets Mamba see local anatomical context before global scanning, and the paper's ablations attribute a 3.2-point mean-PCK gain on MPII to the CMM.
What would settle it
Retrain ViTPose-B and RTMO-L on COCO train2017 without ImageNet pre-training, benchmark all models on the same RTX 4090 with the same input resolution, batch size, and precision, and compare AP and FPS; if ViTPose-B or RTMO-L matches or beats MamKPD-L on both axes, the claimed Mamba advantage does not hold under matched conditions.
Extended reading notes
Core claim
The paper's central claim is that a carefully staged Mamba encoder can deliver accurate 2D keypoint detection without the parameter cost of transformers. Called MamKPD, the network uses a CNN stem, three encoder stages, and a two-deconvolution decoder; each stage pairs a contextual modeling module with 2D selective-scan SS2D blocks. The contextual modeling module uses depth-wise convolutions over the patch grid to capture inter-patch dependencies and a linear layer with a 1x1 depth-wise convolution to distill pose cues inside each patch. The paper reports 77.3% AP on COCO val2017 for MamKPD-L at 1492 FPS, 91.3 mean PCK on MPII, and competitive AP-10K animal pose results, while saving roughly 85% of parameters compared to ViTPose. The authors present the result as evidence that Mamba is a viable backbone class for pose estimation, not just a sequential modeling tool.
Load-bearing premise
The claimed efficiency-and-accuracy advantage assumes all comparison numbers were measured under the same protocol, because MamKPD's speed was measured on an RTX 4090 while some transformer baselines report A100 numbers, and several reproduced baselines were trained without ImageNet pre-training while others used it.
Editorial extensions
If this is right
- A Mamba backbone can serve as the core of a top-down 2D pose estimator, achieving transformer-level COCO AP while using a small fraction of the parameters.
- Real-time pose estimation can run at more than 1400 FPS on a single high-end consumer GPU, according to the paper's measurements.
- Adding local patch-context modeling fixes the main weakness of naive Mamba for images, and the paper shows the gain on MPII.
- The design transfers beyond humans: MamKPD also reports competitive animal keypoint results on AP-10K.
- Because the training uses only MSE heatmap loss with no extra supervision, the architecture is a simple baseline that later Mamba pose methods can build on.
Reading between the lines
- The paper does not explore whether CMM helps other Mamba vision backbones beyond MamKPD; if it does, the same module could be reused in segmentation or detection heads.
- The reported FPS numbers come from an RTX 4090 with no matched hardware benchmark against all baselines, so a head-to-head evaluation with identical GPUs, input resolutions, batch sizes, and precision would settle the speed comparison.
- The parameter savings suggest potential deployment on embedded devices, but the paper only measures a desktop GPU, so edge-latency remains untested.
- The staged Mamba encoder could likely be adapted to single-stage or bottom-up multi-person pose estimation, but those formulations are not explored here.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents MamKPD, a top-down 2D keypoint detector built on a Mamba/SSM backbone. The architecture adds a Context Modeling Module (CMM) before the SS2D blocks in each of three stages, using depthwise convolutions for inter-patch dependencies and linear/1x1 depthwise convolutions for intra-patch feature extraction, followed by a simple deconvolutional decoder. Three model sizes are defined and evaluated on COCO val2017, MPII validation, and AP-10K. The main claims are that MamKPD is the first Mamba-based 2D keypoint detection architecture, that MamKPD-L reaches 77.3 AP on COCO with 1492 FPS on an RTX 4090, that it is state-of-the-art on MPII, and that it is competitive on AP-10K while using far fewer parameters than ViTPose. Ablation studies on MPII examine block counts, channel widths, the stem, and the CMM.
Significance. If the headline comparisons were properly matched, this would be a useful baseline: the model is simple, the parameter counts are low, and the ablation results in Section 4.3 are internally consistent and support the value of the stem and CMM. The claim of being the first Mamba-based 2D keypoint detector is plausible and gives the community a concrete reference architecture. However, the paper ships no code, and the load-bearing speed and SOTA claims are currently supported only by mismatched experimental protocols. The contribution is therefore promising but not yet validated at the level claimed.
major comments (3)
- [Section 4.2.2, Table 2] The speed comparison is not protocol-matched and cannot support the headline FPS claims. MamKPD is a top-down method (Section 4.1 states that an instance detector runs first), so the reported 1492-2564 FPS appears to measure only the keypoint branch on cropped instances. In contrast, RTMO and KAPAO are one-stage methods whose quoted FPS includes person detection and keypoint prediction in a single forward pass. Furthermore, the text notes that ViTPose numbers were measured on an A100 while MamKPD was measured on an RTX 4090, and no input resolution, batch size, numerical precision, or post-processing details are given for any of the speed measurements. Because the '4x faster than ViTPose-L' and real-time advantages rest entirely on these numbers, the authors should re-benchmark all models on the same GPU with a fixed protocol, reporting either end-to-end FPS with a detector or explicitly labeled pose-branch-only FPS against matched pose-branch baselines.
- [Section 4.2.1, Tables 3 and 4] The state-of-the-art claim on MPII is undermined by the dagger footnote. ViTPose-S/B, HRNet-W32, and TokenPose-T are reproduced without ImageNet pre-training, while the published comparison numbers, such as TokenPose-L, use pre-trained backbones. Since MamKPD is also trained from scratch, the 91.3 mean PCK demonstrates superiority only over randomly-initialized re-implementations, not over the published SOTA. The AP-10K comparison in Table 4 has the same issue for the marked baselines. Please use published numbers with the standard pre-training protocol, retrain all baselines under a matched protocol, or explicitly restrict the SOTA claim to the from-scratch setting.
- [Section 4.1, Tables 1 and 2] Several experimental details needed to interpret the efficiency claims are missing: the input resolution used for training and testing, the batch size, the number of warm-up iterations and repeated runs for FPS, whether FP16 or FP32 was used, and whether a flip test or other post-processing was applied at evaluation. The GFLOPs values in Table 1 depend on input resolution, so without this information the parameter-efficiency and speed comparisons cannot be reproduced or verified. These details should be added to the implementation-details paragraph.
minor comments (7)
- [Section 4.2.1, Table 4] HRNet-W48 is listed with 28.5M parameters, which is inconsistent with Table 2 (63.6M) and with the cited paper; please correct the table.
- [Abstract and Section 4.2.2] NVIDIA GTX 4090 should read NVIDIA RTX 4090.
- [Introduction, Section 3.2] The word 'conmmunication' is a typo for 'communication'.
- [Table 3] The Dite-HRNet-18 row is duplicated with different scores; remove the duplicate or label the intended variants.
- [Equation (2) and Figure 2(c)] The notation should clarify whether the SS2D operation is applied once per stage or repeated N_i times per stage, since Figure 2(c) and Table 1 indicate N_i blocks.
- [Section 3.2] The claim that the conventional Mamba module has limited information interaction between patches should be reconciled with the fact that SS2D performs four-directional global scans; otherwise the motivation for CMM is confusing.
- [Abstract] The abstract says MamKPD saves '85% of the parameters compared to ViTPose' without specifying ViTPose-B; please name the exact comparison model, since ViTPose-L would give a different percentage.
Circularity Check
No significant circularity: MamKPD's accuracy and speed results are direct empirical measurements, with no fitted parameter or self-citation chain doing load-bearing work.
full rationale
The paper contains no mathematical derivation chain that could reduce to its own inputs. MamKPD's COCO, MPII, and AP-10K numbers are measured outcomes of trained networks, and the ablation studies validate the Stem and CMM by removal and replacement rather than by fitting a parameter and then re-predicting the same data. The only author self-citations (references [6] and [7]) appear in the introduction as examples of pose-estimation applications and are not load-bearing for the architecture, novelty claim, or experimental results. The FPS comparison in Table 2 is not fully matched: the paper discloses that ViTPose was tested on an A100 while MamKPD was tested on an RTX 4090, and MamKPD is a top-down method whose reported speed appears to exclude the person detector; these are comparison-fairness concerns, not circular reasoning. Similarly, the MPII and AP-10K tables mark reproduced baselines as trained without ImageNet pretraining, which weakens those baselines but does not make any headline claim definitionally equivalent to the paper's inputs. No equation is self-referential, no fitted value is renamed as a prediction, and no external load-bearing claim is imported solely through the authors' own prior work.
Assumptions & free parameters
free parameters (2)
- Stage dimensions =
Small/Base [96,192,384], Large [128,256,512]
- Block counts per stage =
[2,4,6]
assumptions (3)
- domain assumption COCO, MPII, and AP-10K annotations are correct and the official evaluation metrics are meaningful.
- domain assumption The compared baselines are configured equivalently to MamKPD in training and inference.
- domain assumption The SS2D block from VMamba behaves as described in the cited work.
Cite this review
Pith. "Pith review of MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection." pith.science (2026). https://pith.science/paper/2R4BI3WR
@misc{pith2026241201422,
author = {Pith},
title = {Pith review of: MamKPD: A Simple Mamba Baseline for Real-Time 2D Keypoint Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/2R4BI3WR}},
note = {Machine review of arXiv:2412.01422}
}
read the original abstract
Real-time 2D keypoint detection plays an essential role in computer vision. Although CNN-based and Transformer-based methods have achieved breakthrough progress, they often fail to deliver superior performance and real-time speed. This paper introduces MamKPD, the first efficient yet effective mamba-based pose estimation framework for 2D keypoint detection. The conventional Mamba module exhibits limited information interaction between patches. To address this, we propose a lightweight contextual modeling module (CMM) that uses depth-wise convolutions to model inter-patch dependencies and linear layers to distill the pose cues within each patch. Subsequently, by combining Mamba for global modeling across all patches, MamKPD effectively extracts instances' pose information. We conduct extensive experiments on human and animal pose estimation datasets to validate the effectiveness of MamKPD. Our MamKPD-L achieves 77.3% AP on the COCO dataset with 1492 FPS on an NVIDIA GTX 4090 GPU. Moreover, MamKPD achieves state-of-the-art results on the MPII dataset and competitive results on the AP-10K dataset while saving 85% of the parameters compared to ViTPose. Our project page is available at https://mamkpd.github.io/.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Mykhaylo Andriluka, Leonid Pishchulin, Peter V . Gehler, and Bernt Schiele. 2d human pose estimation: New bench- mark and state of the art analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 3686–3693, 2014. 2, 5, 6, 7
work page 2014
-
[2]
Openpose: Realtime multi-person 2d pose es- timation using part affinity fields
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Openpose: Realtime multi-person 2d pose es- timation using part affinity fields. IEEE Trans. Pattern Anal. Mach. Intell., 43(1):172–186, 2021. 1
work page 2021
-
[3]
Human pose estimation with iterative error feedback
Joao Carreira, Pulkit Agrawal, Katerina Fragkiadaki, and Ji- tendra Malik. Human pose estimation with iterative error feedback. In Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 4733– 4742, 2016. 2
work page 2016
-
[4]
Articulated pose estimation by a graphical model with image dependent pairwise rela- tions
Xianjie Chen and Alan L Yuille. Articulated pose estimation by a graphical model with image dependent pairwise rela- tions. In Proceedings of the Neural Information Processing Systems (NeurIPS), 2014. 2
work page 2014
-
[5]
Multi-context attention for hu- man pose estimation
Xiao Chu, Wei Yang, Wanli Ouyang, Cheng Ma, Alan L Yuille, and Xiaogang Wang. Multi-context attention for hu- man pose estimation. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1831–1840, 2017. 2
work page 2017
-
[6]
Relation- based associative joint location for human pose estimation in videos
Yonghao Dang, Jianqin Yin, and Shaojie Zhang. Relation- based associative joint location for human pose estimation in videos. IEEE Trans. Image Process., 31:3973–3986, 2022. 1
work page 2022
-
[7]
Dhrnet: A dual-path hierar- chical relation network for multi-person pose estimation
Yonghao Dang, Jianqin Yin, Liyuan Liu, Pengxiang Ding, Yuan Sun, and Yanzhu Hu. Dhrnet: A dual-path hierar- chical relation network for multi-person pose estimation. Knowledge-Based Systems, 300:112263, 2024. 1
work page 2024
-
[8]
I 2r-net: Intra- and inter-human relation net- work for multi-person pose estimation
Yiwei Ding, Wenjin Deng, Yinglin Zheng, Pengfei Liu, Mei- hong Wang, Xuan Cheng, Jianmin Bao, Dong Chen, and Ming Zeng. I 2r-net: Intra- and inter-human relation net- work for multi-person pose estimation. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pages 855–862, 2022. 6
work page 2022
Show all 62 references
-
[9]
H. Fang, J. Li, H. Tang, C. Xu, H. Zhu, Y . Xiu, Y . Li, and C. Lu. Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time. IEEE Trans. Pattern Anal. Mach. Intell., (1):1–17, 2022. 1
2022
-
[10]
Hungry hungry hippos: To- wards language modeling with state space models
Daniel Y Fu, Tri Dao, Khaled K Saab, Armin W Thomas, Atri Rudra, and Christopher R´e. Hungry hungry hippos: To- wards language modeling with state space models. arXiv preprint arXiv:2212.14052, 2022. 3
2022 arXiv
-
[11]
Mamba: Linear-time sequence mod- eling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence mod- eling with selective state spaces. CoRR, abs/2312.00752,
-
[12]
Hippo: Recurrent memory with optimal polynomial projections
Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christo- pher R´e. Hippo: Recurrent memory with optimal polynomial projections. In Proceedings of the Neural Information Pro- cessing Systems (NeurIPS), pages 1474–1487, 2020. 3
2020
-
[13]
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher R ´e. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021. 3
2021 arXiv
-
[14]
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R ´e. Combining recurrent, convolutional, and continuous-time models with linear state space layers. In Proceedings of the Neural Information Pro- cessing Systems (NeurIPS), pages 572–585, 2021. 1
2021
-
[15]
Bottom-up and top-down reasoning with hierarchical rectified gaussians
Peiyun Hu and Deva Ramanan. Bottom-up and top-down reasoning with hierarchical rectified gaussians. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5600–5609, 2016. 2
2016
-
[16]
LAMP: leveraging language prompts for multi-person pose estimation
Shengnan Hu, Ce Zheng, Zixiang Zhou, Chen Chen, and Gita Sukthankar. LAMP: leveraging language prompts for multi-person pose estimation. In Proceedings of the IEEE International Conference on Intelligent Robots and Systems (IROS), pages 3759–3766, 2023. 2
2023
-
[17]
Localmamba: Visual state space model with windowed selective scan
Tao Huang, Xiaohuan Pei, Shan You, Fei Wang, Chen Qian, and Chang Xu. Localmamba: Visual state space model with windowed selective scan. arXiv preprint arXiv:2403.09338,
-
[18]
Posemamba: Monocular 3d human pose estimation with bidirectional global-local spatio-temporal state space model
Yunlong Huang, Junshuo Liu, Ke Xian, and Robert Caim- ing Qiu. Posemamba: Monocular 3d human pose estimation with bidirectional global-local spatio-temporal state space model. CoRR, abs/2408.03540, 2024. 2, 3
2024 arXiv
-
[19]
Learning human pose estimation features with convolutional networks
Arjun Jain, Jonathan Tompson, Mykhaylo Andriluka, Gra- ham W Taylor, and Christoph Bregler. Learning human pose estimation features with convolutional networks. arXiv preprint arXiv:1312.7302, 2013. 2
2013 arXiv
-
[20]
Rtmpose: Real- time multi-person pose estimation based on mmpose
Tao Jiang, Peng Lu, Li Zhang, Ningsheng Ma, Rui Han, Chengqi Lyu, Yining Li, and Kai Chen. Rtmpose: Real- time multi-person pose estimation based on mmpose. CoRR, abs/2303.07399, 2023. 1, 2, 3
2023 arXiv
-
[21]
Ot- pose: Occlusion-aware transformer for pose estimation in sparsely-labeled videos
Kyung-Min Jin, Gun-Hee Lee, and Seong-Whan Lee. Ot- pose: Occlusion-aware transformer for pose estimation in sparsely-labeled videos. In Proceedings of the IEEE Interna- tional Conference on Systems, Man, and Cybernetics (SMC), pages 3255–3260, 2022. 1
2022
-
[22]
Pose recognition with cascade transformers
Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, and Zhuowen Tu. Pose recognition with cascade transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1944–1953,
1944
-
[23]
Videomamba: State space model for efficient video understanding
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. Videomamba: State space model for efficient video understanding. arXiv preprint arXiv:2403.06977, 2024. 2, 3
2024 arXiv
-
[24]
Dite-hrnet: Dynamic lightweight high-resolution network for human pose estimation
Qun Li, Ziyi Zhang, Fu Xiao, Feng Zhang, and Bir Bhanu. Dite-hrnet: Dynamic lightweight high-resolution network for human pose estimation. In Proceedings of the Inter- national Joint Conference on Artificial Intelligence (IJCAI), pages 1095–1101, 2022. 1, 2, 6, 7
2022
-
[25]
Tokenpose: Learning keypoint tokens for human pose estimation
Yanjie Li, Shoukui Zhang, Zhicheng Wang, Sen Yang, Wankou Yang, Shu-Tao Xia, and Erjin Zhou. Tokenpose: Learning keypoint tokens for human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11293–11302, 2021. 1, 2, 5, 6, 7
2021
-
[26]
Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV), pages 740–755, 2014. 2, 5, 6, 8
2014
-
[27]
Group pose: A sim- ple baseline for end-to-end multi-person pose estimation
Huan Liu, Qiang Chen, Zichang Tan, Jiang-Jiang Liu, Jian Wang, Xiangbo Su, Xiaolong Li, Kun Yao, Junyu Han, Errui Ding, Yao Zhao, and Jingdong Wang. Group pose: A sim- ple baseline for end-to-end multi-person pose estimation. In Proceedings of the IEEE/CVF International Confer...
2023
-
[28]
Swin-umamba: Mamba-based unet with imagenet-based pretraining
Jiarun Liu, Hao Yang, Hong-Yu Zhou, Yan Xi, Lequan Yu, Cheng Li, Yong Liang, Guangming Shi, Yizhou Yu, Shaoting Zhang, Hairong Zheng, and Shanshan Wang. Swin-umamba: Mamba-based unet with imagenet-based pretraining. In Pro- ceedings of the Medical Image Computing and Computer ...
2024
-
[29]
Vmamba: Visual state space model
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, and Yunfan Liu. Vmamba: Visual state space model. CoRR, abs/2401.10166, 2024. 2, 3, 5
2024 arXiv
-
[30]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proceedings of the International Confer- ence on Learning Representations (ICLR), 2019. 5
2019
-
[31]
RTMO: towards high-performance one- stage real-time multi-person pose estimation
Peng Lu, Tao Jiang, Yining Li, Xiangtai Li, Kai Chen, and Wenming Yang. RTMO: towards high-performance one- stage real-time multi-person pose estimation. In Proceed- ings of theIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1491–1500, 2024. 1, 2, 3, 5, 6
2024
-
[32]
Human pose regression by combining indirect part detection and contextual information
Diogo C Luvizon, Hedi Tabia, and David Picard. Human pose regression by combining indirect part detection and contextual information. Computers & Graphics, 85:15–22,
-
[33]
Ppt: Token-pruned pose transformer for monocular and multi-view human pose estimation
Haoyu Ma, Zhe Wang, Yifei Chen, Deying Kong, Liangjian Chen, Xingwei Liu, Xiangyi Yan, Hao Tang, and Xiaohui Xie. Ppt: Token-pruned pose transformer for monocular and multi-view human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV) , pages ...
2022
-
[34]
Tfpose: Direct human pose esti- mation with transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, and Zhibin Wang. Tfpose: Direct human pose esti- mation with transformers. CoRR, 2021. 1, 2
2021
-
[35]
Poseur: Direct human pose regression with transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Xinlong Wang, Zhibin Wang, and Anton van den Hengel. Poseur: Direct human pose regression with transformers. In Pro- ceedings of the European Conference on Computer Vision (ECCV), pages 72–88. Springer, 2022. 2
2022
-
[36]
McNally, Kanav Vats, Alexander Wong, and John McPhee
William J. McNally, Kanav Vats, Alexander Wong, and John McPhee. Rethinking keypoint representations: Modeling keypoints and poses as objects for multi-person human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 37–54, 2022. 1, 2, 6
2022
-
[37]
Efficienthrnet
Christopher Neff, Aneri Sheth, Steven Furgurson, John Mid- dleton, and Hamed Tabkhi. Efficienthrnet. J. Real Time Im- age Process., 18(4):1037–1049, 2021. 1, 2
2021
-
[38]
Stacked hour- glass networks for human pose estimation
Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hour- glass networks for human pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV) , pages 483–499, 2016. 1
2016
-
[39]
Girshick, and Ali Farhadi
Joseph Redmon, Santosh Kumar Divvala, Ross B. Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 779–788, 2016. 1, 2
2016
-
[40]
Bernstein, Alexander C
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition chal- lenge. Int. J. Comput. Vis., 115(3):211–252, 2015. 7
2015
-
[41]
End-to-end multi-person pose estimation with trans- formers
Dahu Shi, Xing Wei, Liangqi Li, Ye Ren, and Wenming Tan. End-to-end multi-person pose estimation with trans- formers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 11069–11078, 2022. 1
2022
-
[42]
Deep high-resolution representation learning for human pose esti- mation
Ke Sun, Bin Xiao, Dong Liu, and Jingdong Wang. Deep high-resolution representation learning for human pose esti- mation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 5693–5703,
-
[43]
Compositional human pose regression
Xiao Sun, Jiaxiang Shang, Shuang Liang, and Yichen Wei. Compositional human pose regression. InProceedings of the IEEE International Conference on Computer Vision (ICCV), pages 2602–2611, 2017. 2
2017
-
[44]
Mingxing Tan and Quoc V . Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In Pro- ceedings of the International Conference on Machine Learn- ing (ICML), pages 6105–6114, 2019. 1, 2
2019
-
[45]
Efficient object localization using convolutional networks
Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, and Christoph Bregler. Efficient object localization using convolutional networks. In Proceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition (CVPR) , pages 648–656, 2015. 2
2015
-
[46]
Joint training of a convolutional network and a graphical model for human pose estimation
Jonathan J Tompson, Arjun Jain, Yann LeCun, and Christoph Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In Proceedings of the Neural Information Processing Systems (NeurIPS) ,
-
[47]
Deeppose: Hu- man pose estimation via deep neural networks
Alexander Toshev and Christian Szegedy. Deeppose: Hu- man pose estimation via deep neural networks. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1653–1660, 2014. 2
2014
-
[48]
Gtpt: Group-based token prun- ing transformer for efficient human pose estimation
Haonan Wang, Jie Liu, Jie Tang, Gangshan Wu, Bo Xu, Yan- bing Chou, and Yong Wang. Gtpt: Group-based token prun- ing transformer for efficient human pose estimation. CoRR, abs/2407.10756, 2024. 2
2024 arXiv
-
[49]
Lite pose: Efficient architecture design for 2d human pose estimation
Yihan Wang, Muyang Li, Han Cai, Wei-Ming Chen, and Song Han. Lite pose: Efficient architecture design for 2d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13126–13136, 2022. 1, 2
2022
-
[50]
Simple baselines for human pose estimation and tracking
Bin Xiao, Haiping Wu, and Yichen Wei. Simple baselines for human pose estimation and tracking. In Proceedings of the European Conference on Computer Vision (ECCV) , pages 472–487, 2018. 1, 2, 3, 5, 6, 7
2018
-
[51]
Swin-pose: Swin transformer based human pose estimation
Zinan Xiong, Chenxi Wang, Ying Li, Yan Luo, and Yu Cao. Swin-pose: Swin transformer based human pose estimation. In Proceedings of the IEEE International Conference on Multimedia Information Processing and Retrieval (MIPR) , pages 228–233, 2022. 1, 6
2022
-
[52]
Vit- pose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vit- pose: Simple vision transformer baselines for human pose estimation. In Proceedings of the Advances in Neural Infor- mation Processing Systems (NeurIPS), 2022. 2, 3, 5, 6, 7
2022
-
[53]
Plainmamba: Improving non-hierarchical mamba in visual recognition
Chenhongyi Yang, Zehui Chen, Miguel Espinosa, Linus Er- icsson, Zhenyu Wang, Jiaming Liu, and Elliot J Crowley. Plainmamba: Improving non-hierarchical mamba in visual recognition. arXiv preprint arXiv:2403.17695, 2024. 3
2024 arXiv
-
[54]
Learning feature pyramids for human pose estimation
Wei Yang, Shuang Li, Wanli Ouyang, Hongsheng Li, and Xiaogang Wang. Learning feature pyramids for human pose estimation. In Proceedings of the IEEE International Confer- ence on Computer Vision (ICCV), pages 1281–1290, 2017. 2
2017
-
[55]
Remamber: Referring image segmentation with mamba twister
Yuhuan Yang, Chaofan Ma, Jiangchao Yao, Zhun Zhong, Ya Zhang, and Yanfeng Wang. Remamber: Referring image segmentation with mamba twister. arXiv preprint arXiv:2403.17839, 2024. 3
2024 arXiv
-
[56]
Lite-hrnet: A lightweight high-resolution network
Changqian Yu, Bin Xiao, Changxin Gao, Lu Yuan, Lei Zhang, Nong Sang, and Jingdong Wang. Lite-hrnet: A lightweight high-resolution network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), pages 10440–10450, 2021. 1, 2, 6, 7
2021
-
[57]
AP-10K: A benchmark for animal pose esti- mation in the wild
Hang Yu, Yufei Xu, Jing Zhang, Wei Zhao, Ziyu Guan, and Dacheng Tao. AP-10K: A benchmark for animal pose esti- mation in the wild. In Proceedings of the Neural Information Processing Systems (NeurIPS), 2021. 2, 5, 7, 8
2021
-
[58]
A survey on visual mamba
Hanwei Zhang, Ying Zhu, Dan Wang, Lijun Zhang, Tianx- iang Chen, and Zi Ye. A survey on visual mamba. CoRR, abs/2404.15956, 2024. 2
2024 arXiv
-
[59]
Efficientpose: Efficient human pose estimation with neural architecture search
Wenqiang Zhang, Jiemin Fang, Xinggang Wang, and Wenyu Liu. Efficientpose: Efficient human pose estimation with neural architecture search. Comput. Vis. Media , 7(3):335– 347, 2021. 1, 2
2021
-
[60]
CLAMP: prompt-based contrastive learn- ing for connecting language and animal pose
Xu Zhang, Wen Wang, Zhe Chen, Yufei Xu, Jing Zhang, and Dacheng Tao. CLAMP: prompt-based contrastive learn- ing for connecting language and animal pose. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 23272–23281, 2023. 2
2023
-
[61]
Deep learning-based human pose estimation: A survey.ACM Comput
Ce Zheng, Wenhan Wu, Chen Chen, Taojiannan Yang, Si- jie Zhu, Ju Shen, Nasser Kehtarnavaz, and Mubarak Shah. Deep learning-based human pose estimation: A survey.ACM Comput. Surv., 56(1):11:1–11:37, 2024. 1
2024
-
[62]
Vision mamba: Efficient visual representation learning with bidirectional state space model
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. Vision mamba: Efficient visual representation learning with bidirectional state space model. In Proceedings of the International Conference on Machine Learning (ICML), 2024. 2
2024
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.