REVIEW 5 major objections 7 minor 1 cited by
XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses
T0 review · 5 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read This paper introduces XRF V2, a multimodal Wi-Fi and wearable IMU dataset, and claims its Mamba-based model XRFMamba localizes 30 daily actions in home settings at 78.74 average mAP while enabling LLM-driven action summarization with…
desk verdict Useful multimodal dataset, but prompt-timestamp ground truth puts a ceiling on what the headline mAP numbers actually mean; worth a serious referee if that issue is addressed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is XRFMamba, whose backbone is a Decomposed Bidirectionally Mamba (DBM) block: a selective state-space model with parameters shared between forward and backward passes, giving linear-time long-range sequence modeling. Before the backbone, a projection layer standardizes channels and a fixed weighted fusion combines Wi-Fi and IMU features with weight 0.2 on Wi-Fi and 0.8 on IMU; the TSSE embedding module, borrowed from WiFiTAD, supplies learned representations in the absence of pretrained sensor models. A feature pyramid prediction head with focal loss for classification and L1 loss for boundary regression produces candidate segments, and temporal non-maximum suppression selects the final detections. The paper's new evaluation object is RMC, which measures whether LLM-generated summaries of predicted action tuples mean the same thing to human auditors as summaries of ground-truth tuples.
What would settle it
Re-annotate a random subset of XRF V2 sequences from the synchronized video with manual start and end labels, then recompute XRFMamba's mAP against those labels; if the manual boundaries differ systematically or the mAP drops substantially below the reported 78.74, the prompt-based annotation is biasing the reported performance.
Extended reading notes
Core claim
The central claim is that temporal action localization and action summarization are feasible from fused Wi-Fi and IMU signals in realistic home settings, not just from video. The paper builds this on XRF V2: 16 volunteers, three environments, 30 actions, 853 valid sequences totaling roughly 16 hours and 16 minutes, with synchronized CSI from one transmitter and three receivers, IMU streams from two phones, two watches, earbuds, and glasses, plus Kinect RGB-D-IR video used for reference and pose extraction. For the localization stage, the paper proposes XRFMamba, which projects and fuses the two modalities with a fixed weighted sum, embeds them with a dual-stream transformer-convolution module, passes them through a decomposed bidirectional Mamba backbone, and predicts action classes and boundaries with a pyramidal head and non-maximum suppression. On the test split, it reports an average mAP of 78.74, beating WiFiTAD by 5.49 mAP points with 35% fewer parameters, and a leave-one-person-out mAP of 72.84. For summarization, the paper defines Response Meaning Consistency (RMC): human auditors compare LLM answers on ground-truth and predicted action tuples for 342 prompt-sequence pairs across three LLM backbones, yielding an average RMC of 0.802.
Load-bearing premise
The ground-truth start and end times are taken directly from the audio prompts that tell volunteers when to start and stop each action, so the annotations assume volunteers switch actions exactly at the prompt boundaries with no reaction delay or transition.
Editorial extensions
If this is right
- Privacy-preserving indoor monitoring can be built from Wi-Fi and IMU signals alone; XRFMamba reaches about 70 mAP zero-shot in an unseen apartment and 72.84 mAP on held-out users.
- Users can trade off device burden and accuracy: a single right-hand smartwatch gives 65.81 mAP, while combining phone, watch, glasses, and ambient Wi-Fi reaches 74.35 mAP, so systems can degrade gracefully when sensors fail.
- Action localization outputs can be converted into useful agent behavior: LLM-based summaries of predicted action tuples match ground-truth summaries in meaning about 80% of the time across medication, drinking, reading, and other health-relevant questions.
- Because XRFMamba uses Mamba's linear-time state-space layers, it processes 30-second windows at about 37 ms latency and an entire action sequence in about 537 ms, making real-time deployment plausible.
Reading between the lines
- Editorial inference: the audio-prompt annotation protocol offers a cheap way to label scripted activities at scale, but it also means any human reaction time shows up as both ground-truth and prediction bias; a small manual re-annotation study on a subset would show how much of the 78.74 mAP is annotation slack.
- Editorial inference: the RMC protocol could be automated end-to-end using the LLM-as-judge consistency prompt the authors already test, which tracks human mRMC closely (0.794 versus 0.802), potentially making action summarization benchmarks cheaper to run.
- Editorial inference: because XRF V2 includes synchronized video-derived 2D pose, it could serve as paired sensor-to-pose and sensor-to-mesh pretraining data for multimodal foundation models, extending beyond localization into action forecasting and synthetic IMU/Wi-Fi data generation.
- Editorial inference: the fixed 0.8 IMU / 0.2 Wi-Fi fusion weight suggests wearable coverage matters more than Wi-Fi link diversity in this setup, but device-ablation experiments are needed to separate sensor placement effects from modality importance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces XRF V2, a multimodal dataset for temporal action localization (TAL) and action summarization in smart-home environments, comprising Wi-Fi CSI, IMU data from phones, watches, earbuds, and glasses, and synchronized Kinect video collected from 16 volunteers across three rooms. The authors propose XRFMamba, a Mamba-based TAL model with a weighted fusion of Wi-Fi and IMU features, and report an average mAP of 78.74 on the dataset, outperforming a re-implemented WiFiTAD by 5.49 points with 35% fewer parameters. They also introduce a new action summarization metric, Response Meaning Consistency (RMC), reporting mRMC of 0.802, along with leave-one-person-out, device-combination, and real-apartment zero-shot experiments. The main empirical claims are supported by ten-run comparisons and paired t-tests, but several load-bearing issues around ground-truth boundary definition, metric definition, baseline re-implementation, and dataset statistics need to be resolved. The dataset and code are publicly released.
Significance. If the validation issues are addressed, XRF V2 would be a useful community resource: it is substantially larger and more diverse than WiFiTAD, covers 30 actions in three environments, provides multiple IMU streams plus Wi-Fi and video, and ships with code and data. The XRFMamba results are strengthened by 10-run t-tests, a device-combination study, a leave-one-person-out evaluation (mAP 72.84), and a real-apartment zero-shot test (mAP around 70), which are valuable beyond the headline number. The proposed RMC metric addresses a genuinely new evaluation need for LLM-based action summarization, although its reliability needs formal inter-annotator validation.
major comments (5)
- [§3.3, Fig. 3] The ground-truth action boundaries are defined as the playback times of audio prompts: 'The start and end times of each audio prompt are automatically recorded as the start and end times of the corresponding action.' Because human reaction and transition latencies are nonzero, every annotated boundary is systematically offset from the actual motor action interval. All methods share the same ground truth, so the relative comparison in Table 4 may remain fair, but the absolute mAP values measure alignment with a prompt schedule rather than with independently verifiable action boundaries. Since synchronized Kinect video is available, the authors should report a video-based boundary validation on a sample of sequences (e.g., manual boundary annotation with mean/median offset), a reaction-delay correction, or an analysis of mAP sensitivity to boundary offsets; without such evidence, the dataset's core annotation validity claim is not established.
- [§6.1, Eq. (8)] The quantity called AP@t is defined as the fraction of ground-truth actions whose tIoU exceeds each threshold, with no ranking, matching, or precision-recall computation. This is a thresholded detection rate, not the average precision used in the video TAL literature from which ActionFormer, TriDet, and TemporalMaxer are drawn. If mAP@avg in Table 4 is obtained by averaging Eq. (8), the headline result is not comparable with published mAP values in that literature. Please clarify the matching and ranking protocol, and either compute standard AP or rename the metric and avoid direct comparisons with published TAL numbers.
- [§6.2.1, Table 4 and Appendix D] The baseline comparison with WiFiTAD is a controlled re-implementation: all methods use the same inputs, embedding, TSSE, and prediction head, with only the backbone replaced. The 'WiFiTAD' row is therefore not the original system of [35], and the reported 5.49-point gain and the paired t-test in Table 12 compare against this variant, not against the published WiFiTAD. If official WiFiTAD cannot be run on multimodal input, this should be stated explicitly and the claim qualified; if the official implementation can be used, the authors should report its numbers or provide evidence that the re-implementation matches the original behavior.
- [§3.4, Table 1, Table 3, and §9] The manuscript contains inconsistent dataset statistics: the abstract and §3.3 say 16 volunteers, while Table 1 lists 15 subjects; §3.4 and Table 3 report 853 valid sequences, while the conclusion says 825. The scene-level counts in Table 3 (305 dining, 320 study, 228 bedroom) also do not match the stated generation protocol of 320/320/240 from 16 volunteers with 5 sequences per scene and 3 or 4 duration variants, and the filtering of 27 sequences is not itemized. Please correct these numbers and explain the filtering procedure, since the dataset statistics are a central contribution.
- [§6.2.2, Table 6, Eq. (9)] RMC is a new metric based on binary human judgments, but the paper reports no inter-annotator agreement statistic. With auditor-level mRMC values ranging from 0.771 to 0.813, chance agreement or systematic auditor bias cannot be ruled out, and the reliability of the reported mRMC=0.802 is under-supported. Report Fleiss' kappa or a comparable agreement measure, state how disagreements were adjudicated, and provide a quantitative human–LLM agreement statistic rather than the qualitative statement that the two are 'closely aligned.'
minor comments (7)
- [§5.1, Eq. (5)] The fusion weight lambda=0.2 is fixed from empirical observation, but no sensitivity analysis is reported; a small sweep of lambda would help establish that the choice is not brittle.
- [§5.2] The 80% truncation rule is reasonable, but it is unclear how a clip boundary that truncates both the head of one action and the tail of the previous action is treated when the remaining segment contains less than 80% of both actions; please clarify the masking and boundary redefinition procedure.
- [§6.2.1, Table 4] The GFlops entry for UWiFiAction is missing; either provide it or explicitly state that it was not reported.
- [§6.2.1, Table 5] The mean row reports mAP@avg = 78.73, whereas the headline and Table 4 report 78.74; please reconcile the rounding or source of the discrepancy.
- [§3.1 and Table 1] Table 1 lists '5 IMUs,' but §3.1 describes two smartwatches, two smartphones, one pair of earbuds, and one pair of smart glasses; please clarify whether the count refers to device types or physical IMU units.
- [§3.1] In the CSI tensor dimension (200t)×1×3×3×30, the trailing ×1 factor is unexplained; a brief note on the single transmitter antenna would remove ambiguity.
- [Eq. (7)] The tIoU formula is correct, but the typesetting makes the denominator ambiguous; adding parentheses around the union computation would improve readability.
Circularity Check
No circularity: the headline mAP is an empirical benchmark against a disclosed annotation protocol, and the RMC metric is defined through independent human/LLM judgment, not through the model's own outputs.
full rationale
The paper's central claims are empirical measurements, not derivations from a first-principles model, and no load-bearing step reduces to its own inputs by construction. The strongest claim, XRFMamba's 78.74 mAP@avg and 5.49-point gain over WiFiTAD, is obtained by training on the XRF V2 training split and evaluating on the held-out test split with standard tIoU-based metrics (Sections 3.4 and 6.2.1). All comparison methods are given the same input, embedding (TSSE), and prediction head, with only the backbone replaced (Section 6.2.1), making the comparison controlled rather than circular. The RMC metric is defined as the fraction of human/LLM judgments that two responses share the same meaning (Eq. 9), and the reported mRMC of 0.802 is a measured agreement rate, not a quantity inferred from the metric's definition. The closest concern is the ground-truth annotation protocol: boundaries are the timestamps of audio prompts played to volunteers (Section 3.3), so reported mAP measures fit to a prompt-derived schedule rather than independently verified human motion. However, this is a data-quality limitation explicitly disclosed in the paper, not a circular derivation: the prompt times are not derived from the sensor streams, the model, or the evaluation metric. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. Self-citations such as XRF55 for hardware setup are background material and are not load-bearing to the performance claims. The paper is therefore self-contained as an empirical benchmark paper, and no circularity step can be exhibited under the required standard.
Assumptions & free parameters
free parameters (2)
- fusion weight lambda =
0.2
- localization loss weight alpha2 =
1000
assumptions (4)
- ad hoc to paper Audio prompt timestamps equal true action start and end times.
- domain assumption Wi-Fi CSI and IMU data from consumer devices contain sufficient discriminative information to distinguish the 30 actions.
- domain assumption Human auditors' consistency judgments are a valid ground truth for action summarization quality.
- domain assumption LLM evaluator consistency scores approximate human judgments.
Cite this review
Pith. "Pith review of XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses." pith.science (2026). https://pith.science/paper/ICD2GRN5
@misc{pith2026250119034,
author = {Pith},
title = {Pith review of: XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses},
year = {2026},
howpublished = {\url{https://pith.science/paper/ICD2GRN5}},
note = {Machine review of arXiv:2501.19034}
}
read the original abstract
Human Action Recognition (HAR) plays a crucial role in applications such as health monitoring, smart home automation, and human-computer interaction. While HAR has been extensively studied, action summarization using Wi-Fi and IMU signals in smart-home environments , which involves identifying and summarizing continuous actions, remains an emerging task. This paper introduces the novel XRF V2 dataset, designed for indoor daily activity Temporal Action Localization (TAL) and action summarization. XRF V2 integrates multimodal data from Wi-Fi signals, IMU sensors (smartphones, smartwatches, headphones, and smart glasses), and synchronized video recordings, offering a diverse collection of indoor activities from 16 volunteers across three distinct environments. To tackle TAL and action summarization, we propose the XRFMamba neural network, which excels at capturing long-term dependencies in untrimmed sensory sequences and achieves the best performance with an average mAP of 78.74, outperforming the recent WiFiTAD by 5.49 points in mAP@avg while using 35% fewer parameters. In action summarization, we introduce a new metric, Response Meaning Consistency (RMC), to evaluate action summarization performance. And it achieves an average Response Meaning Consistency (mRMC) of 0.802. We envision XRF V2 as a valuable resource for advancing research in human action localization, action forecasting, pose estimation, multimodal foundation models pre-training, synthetic data generation, and more. The data and code are available at https://github.com/aiotgroup/XRFV2.
Figures
Figures from the paper (15 more)
Forward citations
Cited by 1 Pith paper
-
Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing
With 10 samples per available activity from a new user, WiAnchor recognizes 10 near-field Wi-Fi activities at 90.4% overall, including 86.3% for the two activities never fine-tuned.
Reference graph
Works this paper leans on
-
[35]
Zhendong Liu, Le Zhang, Bing Li, Yingjie Zhou, Zhenghua Chen, and Ce Zhu. 2025. WiFi CSI Based Temporal Activity Detection Via Dual Pyramid Network. In The 39th Annual AAAI Conference on Artificial Intelligence
work page 2025
-
[1]
Mohammad Arif Ul Alam. 2020. Ai-fairness towards activity recognition of older adults. In 17th EAI International Conference on Mobile and Ubiquitous Systems: Computing, Networking and Services . 108–117
2020
-
[2]
Kamran Ali, Alex X Liu, Wei Wang, and Muhammad Shahzad. 2015. Keystroke recognition using wifi signals. In Proceedings of the 21st annual international conference on mobile computing and networking . 90–102
2015
-
[3]
Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, and Jorge Luis Reyes-Ortiz. 2013. A public domain dataset for human activity recognition using smartphones.. In European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, Vol. 3. 3
2013
-
[4]
Behrooz Azadi, Michael Haslgrübler, Georgios Sopidis, Michaela Murauer, Bernhard Anzengruber, and Alois Ferscha. 2019. Feasibility analysis of unsupervised industrial activity recognition based on a frequent micro action. In Proceedings of the 12th ACM International Conference on PErvasive Technologies Related to Assistive Environments . 368–375
2019
-
[5]
Paramvir Bahl and Venkata N Padmanabhan. 2000. RADAR: An in-building RF-based user location and tracking system. In Proceedings IEEE INFOCOM 2000. Conference on computer communications. Nineteenth annual joint conference of the IEEE computer and communications societies (Cat. No. 00CH37064) , Vol. 2. Ieee, 775–784
2000
-
[6]
Marius Bock, Michael Moeller, and Kristof Van Laerhoven. 2024. Temporal Action Localization for Inertial-based Human Activity Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–19
2024
-
[7]
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition . 961–970
2015
Show all 94 references
-
[8]
Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2019. Openpose: Realtime multi-person 2d pose estimation using part affinity fields. IEEE transactions on pattern analysis and machine intelligence 43, 1 (2019), 172–186
2019
-
[9]
Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang. 2024. Video mamba suite: State space model as a versatile alternative for video understanding. arXiv preprint arXiv:2403.09626 (2024)
2024 arXiv
-
[10]
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2024. A decoder-only foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning
2024
-
[11]
Arne De Brabandere, Jill Emmerzaal, Annick Timmermans, Ilse Jonkers, Benedicte Vanwanseele, and Jesse Davis. 2020. A machine learning approach to estimate hip and knee joint loading using a mobile phone-embedded IMU. Frontiers in bioengineering and biotechnology 8 (2020), 320
2020
-
[12]
Nathan DeVrio, Vimal Mollyn, and Chris Harrison. 2023. SmartPoser: Arm Pose Estimation with a Smartphone and Smartwatch Using UWB and IMU Data. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–11
2023
-
[13]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021
-
[14]
Jiaqi Geng, Dong Huang, and Fernando De la Torre. 2022. Densepose from wifi. arXiv preprint arXiv:2301.00250 (2022)
2022 arXiv
-
[15]
Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)
2023 arXiv
-
[16]
Albert Gu, Karan Goel, and Christopher Ré. 2021. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396 (2021)
2021 arXiv
-
[17]
Daniel Halperin, Wenjun Hu, Anmol Sheth, and David Wetherall. 2011. Tool release: Gathering 802.11 n traces with channel state information. ACM SIGCOMM computer communication review 41, 1 (2011), 53–53. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 0, No. 0, Arti...
2011
-
[18]
Jing He and Wei Yang. 2022. Imar: Multi-user continuous action recognition with wifi signals. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 3 (2022), 1–27
2022
-
[19]
Alexander Hoelzemann, Julia Lee Romero, Marius Bock, Kristof Van Laerhoven, and Qin Lv. 2023. Hang-time HAR: a benchmark dataset for basketball activity recognition using wrist-worn inertial sensors. Sensors 23, 13 (2023), 5879
2023
-
[20]
in the wild
Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. 2017. The thumos challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding 155 (2017), 1–23
2017
-
[21]
Wenjun Jiang, Chenglin Miao, Fenglong Ma, Shuochao Yao, Yaqing Wang, Ye Yuan, Hongfei Xue, Chen Song, Xin Ma, Dimitrios Koutsonikolas, Wenyao Xu, and Lu Su. 2018. Towards environment independent device free human activity recognition. In Proceedings of the 24th annual internat...
2018
-
[22]
Wenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang, Sen Lin, Chong Tian, Srinivasan Murali, Haochen Hu, Zhi Sun, and Lu Su
-
[23]
Woosub Jung, Kenneth Koltermann, Noah Helm, GinaMari Blackwell, Ingrid Pretzer-Aboff, Leslie Cloud, and Gang Zhou. 2022. IMU Sensing Data-Based Kinetic Tremor Detection in Parkinson’s Disease Patients. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Syst...
2022
-
[24]
Manikanta Kotaru, Kiran Joshi, Dinesh Bharadia, and Sachin Katti. 2015. Spotfi: Decimeter level localization using wifi. In Proceedings of the 2015 ACM conference on special interest group on data communication . 269–282
2015
-
[25]
Hilde Kuehne, Ali Arslan, and Thomas Serre. 2014. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the IEEE conference on computer vision and pattern recognition . 780–787
2014
-
[26]
Gierad Laput and Chris Harrison. 2019. Sensing fine-grained hand activity with smartwatches. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–13
2019
-
[27]
Lik-Hang Lee and Pan Hui. 2018. Interaction methods for smart glasses: A survey. IEEE Access 6 (2018), 28712–28732
2018
-
[28]
Hong Li, Wei Yang, Jianxin Wang, Yang Xu, and Liusheng Huang. 2016. WiFinger: Talk to your smart devices with finger-grained gesture. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing . 250–261
2016
-
[29]
Shengjie Li, Zhaopeng Liu, Qin Lv, Yanyan Zou, Yue Zhang, and Daqing Zhang. 2024. WiLife: Long-term Daily Status Monitoring and Habit Mining of the Elderly Leveraging Ubiquitous Wi-Fi Signals. ACM Transactions on Computing for Healthcare (2024)
2024
-
[30]
Xiang Li, Shengjie Li, Daqing Zhang, Jie Xiong, Yasha Wang, and Hong Mei. 2016. Dynamic-MUSIC: Accurate device-free indoor localization. In Proceedings of the 2016 ACM international joint conference on pervasive and ubiquitous computing . 196–207
2016
-
[31]
Jensen, and Bin Yang
Zhe Li, Xiangfei Qiu, Peng Chen, Yihang Wang, Hanyin Cheng, Yang Shu, Jilin Hu, Chenjuan Guo, Aoying Zhou, Qingsong Wen, Christian S. Jensen, and Bin Yang. 2024. Foundts: Comprehensive and unified benchmarking of foundation models for time series forecasting. arXiv preprint ar...
2024 arXiv
-
[32]
Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2117–2125
2017
-
[33]
Girshick, Kaiming He, and Piotr Dollar
Tsung-Yi Lin, Piotr Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. 2017. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 2980–2988
2017
-
[34]
Yi Liu, Limin Wang, Yali Wang, Xiao Ma, and Yu Qiao. 2022. Fineaction: A fine-grained video dataset for temporal action localization. IEEE transactions on image processing 31 (2022), 6937–6950
2022
-
[36]
Yongsen Ma, Gang Zhou, Shuangquan Wang, Hongyang Zhao, and Woosub Jung. 2018. SignFi: Sign language recognition using WiFi. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 1 (2018), 1–21
2018
-
[37]
Huina Meng, Xilei Wu, Xin Wang, Yuhan Fan, Jingang Shi, Han Ding, and Fei Wang. 2022. Mask wearing status estimation with smartwatches. arXiv preprint arXiv:2205.06113 (2022)
2022 arXiv
-
[38]
Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison, and Karan Ahuja. 2023. Imuposer: Full-body pose estimation using imus in phones, watches, and earbuds. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–12
2023
-
[39]
Jay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh, and Romit Roy Choudhury. 2020. EarSense: earphones as a teeth activity sensor. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking . 1–13
2020
-
[40]
Qifan Pu, Sidhant Gupta, Shyamnath Gollakota, and Shwetak Patel. 2013. Whole-home gesture recognition using wireless signals. In Proceedings of the 19th annual international conference on Mobile computing & networking . 27–38
2013
-
[41]
Philipp M Scholl, Matthias Wille, and Kristof Van Laerhoven. 2015. Wearables in the wet lab: a laboratory system for capturing and guiding experiments. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing . 589–599
2015
-
[42]
Meng Shang, Lenore Dedeyne, Jolan Dupont, Laura Vercauteren, Nadjia Amini, Laurence Lapauw, Evelien Gielen, Sabine Verschueren, Carolina Varon, Walter De Raedt, and Bart Vanrumste. 2024. Otago exercises monitoring for older adults by a single imu and hierarchical machine learn...
2024
-
[43]
Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Li, and Dacheng Tao. 2023. Tridet: Temporal action detection with relative boundary modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18857–18866
2023
-
[44]
Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, and Ming Jin. 2024. Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts. arXiv preprint arXiv:2409.16040 (2024)
2024 arXiv
-
[45]
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016. Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1049–1058
2016
-
[46]
Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. 2022. Simplified state space layers for sequence modeling. arXiv preprint arXiv:2208.04933 (2022)
2022 arXiv
-
[47]
Georgios Sopidis, Michael Haslgrübler, Behrooze Azadi, Bernhard Anzengruber-Tánase, Abdelrahman Ahmad, Alois Ferscha, and Martin Baresch. 2022. Micro-activity recognition in industrial assembly process with IMU data and deep learning. In Proceedings of the 15th International C...
2022
-
[48]
Sebastian Stein and Stephen J McKenna. 2013. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In Proceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing . 729–738
2013
-
[49]
Ke Sun, Chunyu Xia, Xinyu Zhang, Hao Chen, and Charlie Jianzhong Zhang. 2024. Multimodal Daily-Life Logging in Free-living Environment Using Non-Visual Egocentric Sensors on a Smartphone. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 ...
2024
-
[50]
Sheng Tan and Jie Yang. 2016. WiFinger: Leveraging commodity WiFi for fine-grained finger gesture recognition. In Proceedings of the 17th ACM international symposium on mobile ad hoc networking and computing . 201–210
2016
-
[51]
Tuan N Tang, Kwonyoung Kim, and Kwanghoon Sohn. 2023. Temporalmaxer: Maximize temporal context with only max pooling for temporal action localization. arXiv preprint arXiv:2303.09055 (2023)
2023 arXiv
-
[52]
Jiacheng Tian, Pan Zhou, Fangmin Sun, Tao Wang, and Hexiang Zhang. 2021. Wearable IMU-based gym exercise recognition using data fusion methods. In The Fifth International Conference on Biological Information and Biomedical Engineering . 1–7
2021
-
[53]
Yonglong Tian, Guang-He Lee, Hao He, Chen-Yu Hsu, and Dina Katabi. 2018. RF-based fall monitoring using convolutional neural networks. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 3 (2018), 1–24
2018
-
[54]
Dhruv Verma, Sejal Bhalla, Dhruv Sahnan, Jainendra Shukla, and Aman Parnami. 2021. Expressear: Sensing fine-grained facial expressions with earables. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 3 (2021), 1–28
2021
-
[55]
Florian Wahl, Martin Freund, and Oliver Amft. 2015. Using smart eyeglasses as a wearable game controller. In Adjunct Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2015 ACM International Symposium on Wear...
2015
-
[56]
Fei Wang, Jianwei Feng, Yinliang Zhao, Xiaobin Zhang, Shiyuan Zhang, and Jinsong Han. 2019. Joint Activity Recognition and Indoor Localization With WiFi Fingerprints. IEEE Access 7 (2019), 80058–80068
2019
-
[57]
Fei Wang, Yiao Gao, Bo Lan, Han Ding, Jingang Shi, and Jinsong Han. 2023. U-Shape Networks Are Unified Backbones for Human Action Understanding From Wi-Fi Signals. IEEE Internet of Things Journal (2023)
2023
-
[58]
Fei Wang, Yizhe Lv, Mengdie Zhu, Han Ding, and Jinsong Han. 2024. XRF55: A Radio Frequency Dataset for Human Indoor Action Analysis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 (2024), 1–34
2024
-
[59]
Fei Wang, Stanislav Panev, Ziyi Dai, Jinsong Han, and Dong Huang. 2019. Can WiFi estimate person pose?arXiv preprint arXiv:1904.00277 (2019)
2019 arXiv
-
[60]
Fei Wang, Tingting Zhang, Xilei Wu, Pengcheng Wang, Xin Wang, Han Ding, Jingang Shi, Jinsong Han, and Dong Huang. 2025. You Can Wash Hands Better: Accurate Daily Handwashing Assessment with a Smartwatch. IEEE Transactions on Mobile Computing (2025)
2025
-
[61]
Fei Wang, Sanping Zhou, Stanislav Panev, Jinsong Han, and Dong Huang. 2019. Person-in-WiFi: Fine-grained person perception using WiFi. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5452–5461
2019
-
[62]
Hao Wang, Daqing Zhang, Junyi Ma, Yasha Wang, Yuxiang Wang, Dan Wu, Tao Gu, and Bing Xie. 2016. Human respiration detection with commodity WiFi devices: Do user location and body orientation matter?. InProceedings of the 2016 ACM international joint conference on pervasive and...
2016
-
[63]
Wei Wang, Alex X Liu, Muhammad Shahzad, Kang Ling, and Sanglu Lu. 2015. Understanding and modeling of wifi signal based human activity recognition. In Proceedings of the 21st annual international conference on mobile computing and networking . 65–76
2015
-
[64]
Wei Wang, Alex X Liu, Muhammad Shahzad, Kang Ling, and Sanglu Lu. 2017. Device-free human activity recognition using commercial WiFi devices. IEEE Journal on Selected Areas in Communications 35, 5 (2017), 1118–1131
2017
-
[65]
Xin Wang, Xilei Wu, Huina Meng, Yuhan Fan, Jingang Shi, Han Ding, and Fei Wang. 2022. Social distancing alert with smartwatches. arXiv preprint arXiv:2205.06110 (2022)
2022 arXiv
-
[66]
Xuyu Wang, Chao Yang, and Shiwen Mao. 2017. TensorBeat: Tensor decomposition for monitoring multiperson breathing beats with commodity WiFi. ACM Transactions on Intelligent Systems and Technology (TIST) 9, 1 (2017), 1–27
2017
-
[67]
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, Limin Wang, and Yu Qiao. 2022. Internvideo: General video foundation models via generative and discriminativ...
2022 arXiv
-
[68]
Yan Wang, Jian Liu, Yingying Chen, Marco Gruteser, Jie Yang, and Hongbo Liu. 2014. E-eyes: Device-free location-oriented activity identification using fine-grained WiFi signatures. In Proceedings of the 20th annual international conference on Mobile computing and networking. 617–628
2014
-
[69]
Yichao Wang, Yili Ren, Yingying Chen, and Jie Yang. 2022. Wi-mesh: A wifi vision-based approach for 3d human mesh construction. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems . 362–376
2022
-
[70]
Yichao Wang, Yili Ren, Yingying Chen, and Jie Yang. 2022. A wifi vision-based 3D human mesh reconstruction. In Proceedings of the 28th Annual International Conference on Mobile Computing and Networking . 814–816
2022
-
[71]
Yuxi Wang, Kaishun Wu, and Lionel M Ni. 2016. Wifall: Device-free fall detection by wireless networks. IEEE Transactions on Mobile Computing 16, 2 (2016), 581–594
2016
-
[72]
Johann P Wolff, Florian Grützmacher, Arne Wellnitz, and Christian Haubelt. 2018. Activity recognition using head worn inertial sensors. In Proceedings of the 5th international Workshop on Sensor-based Activity Recognition and Interaction . 1–7
2018
-
[73]
Chenshu Wu, Zheng Yang, Yunhao Liu, and Wei Xi. 2012. WILL: Wireless indoor localization without site survey. IEEE Transactions on Parallel and Distributed systems 24, 4 (2012), 839–848
2012
-
[74]
Kaishun Wu, Haoyu Tan, Hoilun Ngan, Yunhuai Liu, and Lionel M Ni. 2011. Chip error pattern analysis in IEEE 802.15. 4. IEEE Transactions on Mobile Computing 11, 4 (2011), 543–552
2011
-
[75]
Kaishun Wu, Jiang Xiao, Youwen Yi, Min Gao, and Lionel M Ni. 2012. FILA: Fine-grained indoor localization. In 2012 Proceedings IEEE INFOCOM. IEEE, 2210–2218
2012
-
[76]
Wei Xi, Jizhong Zhao, Xiang-Yang Li, Kun Zhao, Shaojie Tang, Xue Liu, and Zhiping Jiang. 2014. Electronic frog eye: Counting crowd using WiFi. In IEEE INFOCOM 2014-IEEE Conference on Computer Communications . IEEE, 361–369
2014
-
[77]
Rui Xiao, Jianwei Liu, Jinsong Han, and Kui Ren. 2021. OneFi: One-Shot Recognition for Unseen Gesture via COTS WiFi. In Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems . 206–219
2021
-
[78]
Wentao Xie, Qian Zhang, and Jin Zhang. 2021. Acoustic-based upper facial action recognition for smart eyewear. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 2 (2021), 1–28
2021
-
[79]
Yaxiong Xie, Zhenjiang Li, and Mo Li. 2015. Precise power delay profiling with commodity WiFi. In Proceedings of the 21st Annual international conference on Mobile Computing and Networking . 53–64
2015
-
[80]
Leiyang Xu, Xiaolong Zheng, Xinrun Du, Liang Liu, and Huadong Ma. 2024. WiCamera: Vortex Electromagnetic Wave-Based WiFi Imaging. IEEE Transactions on Mobile Computing (2024)
2024
-
[81]
Kangwei Yan, Fei Wang, Bo Qian, Han Ding, Jinsong Han, and Xing Wei. 2024. Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-Fi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 969–978
2024
-
[82]
Jianfei Yang, He Huang, Yunjiao Zhou, Xinyan Chen, Yuecong Xu, Shenghai Yuan, Han Zou, Chris Xiaoxuan Lu, and Lihua Xie. 2024. Mm-fi: Multi-modal non-intrusive 4d human dataset for versatile wireless sensing. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[83]
Zheng Yang, Zimu Zhou, and Yunhao Liu. 2013. From RSSI to CSI: Indoor localization via channel response. ACM Computing Surveys (CSUR) 46, 2 (2013), 1–32
2013
-
[84]
Bohan Yu, Yuxiang Wang, Kai Niu, Youwei Zeng, Tao Gu, Leye Wang, Cuntai Guan, and Daqing Zhang. 2021. WiFi-sleep: Sleep stage monitoring using commodity Wi-Fi devices. IEEE internet of things journal 8, 18 (2021), 13900–13913
2021
-
[85]
Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022. Actionformer: Localizing moments of actions with transformers. In European Conference on Computer Vision. Springer, 492–510
2022
-
[86]
Jie Zhang, Zhanyong Tang, Meng Li, Dingyi Fang, Petteri Nurmi, and Zheng Wang. 2018. CrossSense: Towards cross-site and large-scale WiFi sensing. In Proceedings of the 24th annual international conference on mobile computing and networking . 305–320
2018
-
[87]
Guangrong Zhao, Yiran Shen, Feng Li, Lei Liu, Lizhen Cui, and Hongkai Wen. 2024. Ui-Ear: On-face Gesture Recognition Through On-ear Vibration Sensing. IEEE Transactions on Mobile Computing (2024)
2024
-
[88]
Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. 2019. Hacs: Human action clips and segments dataset for recognition and temporal localization. In Proceedings of the IEEE International Conference on Computer Vision . 8668–8678
2019
-
[89]
SHENGDONG ZHAO, FELICIA TAN, and KATHERINE FENNEDY. 2023. Heads-Up Computing. Commun. ACM 66, 9 (2023)
2023
-
[90]
Xiaolong Zheng, Jiliang Wang, Longfei Shangguan, Zimu Zhou, and Yunhao Liu. 2016. Smokey: Ubiquitous smoking detection with commercial WiFi infrastructures. In IEEE INFOCOM 2016-The 35th Annual IEEE International Conference on Computer Communications . IEEE, 1–9
2016
-
[91]
Yue Zheng, Yi Zhang, Kun Qian, Guidong Zhang, Yunhao Liu, Chenshu Wu, and Zheng Yang. 2019. Zero-effort cross-domain gesture recognition with Wi-Fi. In Proceedings of the 17th annual international conference on mobile systems, applications, and services . 313–325
2019
-
[92]
Luowei Zhou, Chenliang Xu, and Jason Corso. 2018. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32
2018
-
[93]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417 (2024). Proc. ACM Interact. Mob. Wearable Ubiquitous Technol....
2024 arXiv
-
[2020]
In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking
Towards 3D human pose construction using WiFi. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking. 1–14
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.