Pith. sign in

REVIEW 5 major objections 7 minor 1 cited by

XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses

T0 review · 5 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This paper introduces XRF V2, a multimodal Wi-Fi and wearable IMU dataset, and claims its Mamba-based model XRFMamba localizes 30 daily actions in home settings at 78.74 average mAP while enabling LLM-driven action summarization with…

desk verdict Useful multimodal dataset, but prompt-timestamp ground truth puts a ceiling on what the headline mAP numbers actually mean; worth a serious referee if that issue is addressed. read the letter →

arxiv 2501.19034 v2 pith:ICD2GRN5 submitted 2025-01-31 cs.CV

classification cs.CV
keywords temporalactionlocalizationsummarizationWi-FisensingIMUMambasmarthomemultimodaldatasetwearabledevices
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that continuous indoor daily activities can be located in time and summarized in plain language using only ubiquitous sensors—Wi-Fi channel state information plus IMUs in phones, watches, earbuds, and glasses—without relying on cameras. To make this case, it contributes XRF V2, a dataset of 853 annotated action sequences from 16 volunteers in three rooms, spanning 30 everyday actions, and a model, XRFMamba, built on Mamba state-space layers. The paper reports that XRFMamba outperforms prior Wi-Fi temporal action localization methods, reaching 78.74 average mAP versus 73.25 for WiFiTAD while using 35% fewer parameters, and that feeding its predicted action tuples to LLMs yields summaries judged meaning-consistent with ground truth 0.802 of the time. If the claims hold, privacy-preserving smart-home assistants, health monitors, and elderly-care systems could operate on sensors already present in people's pockets and on their wrists.

What carries the argument

The load-bearing object is XRFMamba, whose backbone is a Decomposed Bidirectionally Mamba (DBM) block: a selective state-space model with parameters shared between forward and backward passes, giving linear-time long-range sequence modeling. Before the backbone, a projection layer standardizes channels and a fixed weighted fusion combines Wi-Fi and IMU features with weight 0.2 on Wi-Fi and 0.8 on IMU; the TSSE embedding module, borrowed from WiFiTAD, supplies learned representations in the absence of pretrained sensor models. A feature pyramid prediction head with focal loss for classification and L1 loss for boundary regression produces candidate segments, and temporal non-maximum suppression selects the final detections. The paper's new evaluation object is RMC, which measures whether LLM-generated summaries of predicted action tuples mean the same thing to human auditors as summaries of ground-truth tuples.

What would settle it

Re-annotate a random subset of XRF V2 sequences from the synchronized video with manual start and end labels, then recompute XRFMamba's mAP against those labels; if the manual boundaries differ systematically or the mAP drops substantially below the reported 78.74, the prompt-based annotation is biasing the reported performance.

Watch

Extended reading notes

Core claim

The central claim is that temporal action localization and action summarization are feasible from fused Wi-Fi and IMU signals in realistic home settings, not just from video. The paper builds this on XRF V2: 16 volunteers, three environments, 30 actions, 853 valid sequences totaling roughly 16 hours and 16 minutes, with synchronized CSI from one transmitter and three receivers, IMU streams from two phones, two watches, earbuds, and glasses, plus Kinect RGB-D-IR video used for reference and pose extraction. For the localization stage, the paper proposes XRFMamba, which projects and fuses the two modalities with a fixed weighted sum, embeds them with a dual-stream transformer-convolution module, passes them through a decomposed bidirectional Mamba backbone, and predicts action classes and boundaries with a pyramidal head and non-maximum suppression. On the test split, it reports an average mAP of 78.74, beating WiFiTAD by 5.49 mAP points with 35% fewer parameters, and a leave-one-person-out mAP of 72.84. For summarization, the paper defines Response Meaning Consistency (RMC): human auditors compare LLM answers on ground-truth and predicted action tuples for 342 prompt-sequence pairs across three LLM backbones, yielding an average RMC of 0.802.

Load-bearing premise

The ground-truth start and end times are taken directly from the audio prompts that tell volunteers when to start and stop each action, so the annotations assume volunteers switch actions exactly at the prompt boundaries with no reaction delay or transition.

Editorial extensions

If this is right

  • Privacy-preserving indoor monitoring can be built from Wi-Fi and IMU signals alone; XRFMamba reaches about 70 mAP zero-shot in an unseen apartment and 72.84 mAP on held-out users.
  • Users can trade off device burden and accuracy: a single right-hand smartwatch gives 65.81 mAP, while combining phone, watch, glasses, and ambient Wi-Fi reaches 74.35 mAP, so systems can degrade gracefully when sensors fail.
  • Action localization outputs can be converted into useful agent behavior: LLM-based summaries of predicted action tuples match ground-truth summaries in meaning about 80% of the time across medication, drinking, reading, and other health-relevant questions.
  • Because XRFMamba uses Mamba's linear-time state-space layers, it processes 30-second windows at about 37 ms latency and an entire action sequence in about 537 ms, making real-time deployment plausible.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the audio-prompt annotation protocol offers a cheap way to label scripted activities at scale, but it also means any human reaction time shows up as both ground-truth and prediction bias; a small manual re-annotation study on a subset would show how much of the 78.74 mAP is annotation slack.
  • Editorial inference: the RMC protocol could be automated end-to-end using the LLM-as-judge consistency prompt the authors already test, which tracks human mRMC closely (0.794 versus 0.802), potentially making action summarization benchmarks cheaper to run.
  • Editorial inference: because XRF V2 includes synchronized video-derived 2D pose, it could serve as paired sensor-to-pose and sensor-to-mesh pretraining data for multimodal foundation models, extending beyond localization into action forecasting and synthetic IMU/Wi-Fi data generation.
  • Editorial inference: the fixed 0.8 IMU / 0.2 Wi-Fi fusion weight suggests wearable coverage matters more than Wi-Fi link diversity in this setup, but device-ablation experiments are needed to separate sensor placement effects from modality importance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper introduces XRF V2, a multimodal dataset for temporal action localization (TAL) and action summarization in smart-home environments, comprising Wi-Fi CSI, IMU data from phones, watches, earbuds, and glasses, and synchronized Kinect video collected from 16 volunteers across three rooms. The authors propose XRFMamba, a Mamba-based TAL model with a weighted fusion of Wi-Fi and IMU features, and report an average mAP of 78.74 on the dataset, outperforming a re-implemented WiFiTAD by 5.49 points with 35% fewer parameters. They also introduce a new action summarization metric, Response Meaning Consistency (RMC), reporting mRMC of 0.802, along with leave-one-person-out, device-combination, and real-apartment zero-shot experiments. The main empirical claims are supported by ten-run comparisons and paired t-tests, but several load-bearing issues around ground-truth boundary definition, metric definition, baseline re-implementation, and dataset statistics need to be resolved. The dataset and code are publicly released.

Significance. If the validation issues are addressed, XRF V2 would be a useful community resource: it is substantially larger and more diverse than WiFiTAD, covers 30 actions in three environments, provides multiple IMU streams plus Wi-Fi and video, and ships with code and data. The XRFMamba results are strengthened by 10-run t-tests, a device-combination study, a leave-one-person-out evaluation (mAP 72.84), and a real-apartment zero-shot test (mAP around 70), which are valuable beyond the headline number. The proposed RMC metric addresses a genuinely new evaluation need for LLM-based action summarization, although its reliability needs formal inter-annotator validation.

major comments (5)
  1. [§3.3, Fig. 3] The ground-truth action boundaries are defined as the playback times of audio prompts: 'The start and end times of each audio prompt are automatically recorded as the start and end times of the corresponding action.' Because human reaction and transition latencies are nonzero, every annotated boundary is systematically offset from the actual motor action interval. All methods share the same ground truth, so the relative comparison in Table 4 may remain fair, but the absolute mAP values measure alignment with a prompt schedule rather than with independently verifiable action boundaries. Since synchronized Kinect video is available, the authors should report a video-based boundary validation on a sample of sequences (e.g., manual boundary annotation with mean/median offset), a reaction-delay correction, or an analysis of mAP sensitivity to boundary offsets; without such evidence, the dataset's core annotation validity claim is not established.
  2. [§6.1, Eq. (8)] The quantity called AP@t is defined as the fraction of ground-truth actions whose tIoU exceeds each threshold, with no ranking, matching, or precision-recall computation. This is a thresholded detection rate, not the average precision used in the video TAL literature from which ActionFormer, TriDet, and TemporalMaxer are drawn. If mAP@avg in Table 4 is obtained by averaging Eq. (8), the headline result is not comparable with published mAP values in that literature. Please clarify the matching and ranking protocol, and either compute standard AP or rename the metric and avoid direct comparisons with published TAL numbers.
  3. [§6.2.1, Table 4 and Appendix D] The baseline comparison with WiFiTAD is a controlled re-implementation: all methods use the same inputs, embedding, TSSE, and prediction head, with only the backbone replaced. The 'WiFiTAD' row is therefore not the original system of [35], and the reported 5.49-point gain and the paired t-test in Table 12 compare against this variant, not against the published WiFiTAD. If official WiFiTAD cannot be run on multimodal input, this should be stated explicitly and the claim qualified; if the official implementation can be used, the authors should report its numbers or provide evidence that the re-implementation matches the original behavior.
  4. [§3.4, Table 1, Table 3, and §9] The manuscript contains inconsistent dataset statistics: the abstract and §3.3 say 16 volunteers, while Table 1 lists 15 subjects; §3.4 and Table 3 report 853 valid sequences, while the conclusion says 825. The scene-level counts in Table 3 (305 dining, 320 study, 228 bedroom) also do not match the stated generation protocol of 320/320/240 from 16 volunteers with 5 sequences per scene and 3 or 4 duration variants, and the filtering of 27 sequences is not itemized. Please correct these numbers and explain the filtering procedure, since the dataset statistics are a central contribution.
  5. [§6.2.2, Table 6, Eq. (9)] RMC is a new metric based on binary human judgments, but the paper reports no inter-annotator agreement statistic. With auditor-level mRMC values ranging from 0.771 to 0.813, chance agreement or systematic auditor bias cannot be ruled out, and the reliability of the reported mRMC=0.802 is under-supported. Report Fleiss' kappa or a comparable agreement measure, state how disagreements were adjudicated, and provide a quantitative human–LLM agreement statistic rather than the qualitative statement that the two are 'closely aligned.'
minor comments (7)
  1. [§5.1, Eq. (5)] The fusion weight lambda=0.2 is fixed from empirical observation, but no sensitivity analysis is reported; a small sweep of lambda would help establish that the choice is not brittle.
  2. [§5.2] The 80% truncation rule is reasonable, but it is unclear how a clip boundary that truncates both the head of one action and the tail of the previous action is treated when the remaining segment contains less than 80% of both actions; please clarify the masking and boundary redefinition procedure.
  3. [§6.2.1, Table 4] The GFlops entry for UWiFiAction is missing; either provide it or explicitly state that it was not reported.
  4. [§6.2.1, Table 5] The mean row reports mAP@avg = 78.73, whereas the headline and Table 4 report 78.74; please reconcile the rounding or source of the discrepancy.
  5. [§3.1 and Table 1] Table 1 lists '5 IMUs,' but §3.1 describes two smartwatches, two smartphones, one pair of earbuds, and one pair of smart glasses; please clarify whether the count refers to device types or physical IMU units.
  6. [§3.1] In the CSI tensor dimension (200t)×1×3×3×30, the trailing ×1 factor is unexplained; a brief note on the single transmitter antenna would remove ambiguity.
  7. [Eq. (7)] The tIoU formula is correct, but the typesetting makes the denominator ambiguous; adding parentheses around the union computation would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the headline mAP is an empirical benchmark against a disclosed annotation protocol, and the RMC metric is defined through independent human/LLM judgment, not through the model's own outputs.

full rationale

The paper's central claims are empirical measurements, not derivations from a first-principles model, and no load-bearing step reduces to its own inputs by construction. The strongest claim, XRFMamba's 78.74 mAP@avg and 5.49-point gain over WiFiTAD, is obtained by training on the XRF V2 training split and evaluating on the held-out test split with standard tIoU-based metrics (Sections 3.4 and 6.2.1). All comparison methods are given the same input, embedding (TSSE), and prediction head, with only the backbone replaced (Section 6.2.1), making the comparison controlled rather than circular. The RMC metric is defined as the fraction of human/LLM judgments that two responses share the same meaning (Eq. 9), and the reported mRMC of 0.802 is a measured agreement rate, not a quantity inferred from the metric's definition. The closest concern is the ground-truth annotation protocol: boundaries are the timestamps of audio prompts played to volunteers (Section 3.3), so reported mAP measures fit to a prompt-derived schedule rather than independently verified human motion. However, this is a data-quality limitation explicitly disclosed in the paper, not a circular derivation: the prompt times are not derived from the sensor streams, the model, or the evaluation metric. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. Self-citations such as XRF55 for hardware setup are background material and are not load-bearing to the performance claims. The paper is therefore self-contained as an empirical benchmark paper, and no circularity step can be exhibited under the required standard.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claims rest on empirical benchmark results rather than derivations. The main free choices are the fusion weight and loss balance hyperparameters. The most significant assumptions are about annotation validity and the sufficiency of the sensing modalities.

free parameters (2)
  • fusion weight lambda = 0.2
    Weight for Wi-Fi branch in the weighted fusion, chosen because IMU was observed to perform better (Section 5.1).
  • localization loss weight alpha2 = 1000
    Balance term for the localization loss relative to classification loss, set empirically (Section 5.1, Eq. 6).
assumptions (4)
  • ad hoc to paper Audio prompt timestamps equal true action start and end times.
    Section 3.3 states start and end times are recorded from prompts, assuming volunteers execute actions exactly at prompt boundaries.
  • domain assumption Wi-Fi CSI and IMU data from consumer devices contain sufficient discriminative information to distinguish the 30 actions.
    This is the foundational sensing assumption for the dataset and model evaluation.
  • domain assumption Human auditors' consistency judgments are a valid ground truth for action summarization quality.
    Section 6.2.2 uses five auditors without inter-rater reliability analysis and treats their majority as ground truth.
  • domain assumption LLM evaluator consistency scores approximate human judgments.
    Section 6.2.2 reports mRMC(LLMs)=0.794 versus mRMC(humans)=0.802, but does not establish equivalence beyond the aggregate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses." pith.science (2026). https://pith.science/paper/ICD2GRN5

@misc{pith2026250119034,
  author       = {Pith},
  title        = {Pith review of: XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ICD2GRN5}},
  note         = {Machine review of arXiv:2501.19034}
}
read the original abstract

Human Action Recognition (HAR) plays a crucial role in applications such as health monitoring, smart home automation, and human-computer interaction. While HAR has been extensively studied, action summarization using Wi-Fi and IMU signals in smart-home environments , which involves identifying and summarizing continuous actions, remains an emerging task. This paper introduces the novel XRF V2 dataset, designed for indoor daily activity Temporal Action Localization (TAL) and action summarization. XRF V2 integrates multimodal data from Wi-Fi signals, IMU sensors (smartphones, smartwatches, headphones, and smart glasses), and synchronized video recordings, offering a diverse collection of indoor activities from 16 volunteers across three distinct environments. To tackle TAL and action summarization, we propose the XRFMamba neural network, which excels at capturing long-term dependencies in untrimmed sensory sequences and achieves the best performance with an average mAP of 78.74, outperforming the recent WiFiTAD by 5.49 points in mAP@avg while using 35% fewer parameters. In action summarization, we introduce a new metric, Response Meaning Consistency (RMC), to evaluate action summarization performance. And it achieves an average Response Meaning Consistency (mRMC) of 0.802. We envision XRF V2 as a valuable resource for advancing research in human action localization, action forecasting, pose estimation, multimodal foundation models pre-training, synthetic data generation, and more. The data and code are available at https://github.com/aiotgroup/XRFV2.

Figures

Figures reproduced from arXiv: 2501.19034 by the authors.

Figure 1
Figure 1. XRF V2 includes multimodal data from Wi-Fi signals, IMU sensors (smartphones, smartwatches, headphones, and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the experimental scenarios, where volunteers perform continuous action sequences in the dining room, [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Automatic annotation process. The audio prompts are played in real-time to guide the volunteers. Upon hearing an [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Examples of action sequence from the study room, dining room, and bedroom, with a detailed visualization of the [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: State space model (a). The model can be computed in linear recurrence (b) for inference, or global convolution (c) for [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Basic Mamba Block [15]. Due to the flexibility of SSMs, they can switch to convolution during training and to recurrence during inference, providing fast performance in both stages. SSMs have a rich foundational knowledge, and introducing all aspects is beyond the scop…
Figure 7
Figure 7. Figure 7: XRFMamba uses Mamba as the feature learning backbone. It takes an input sequence and predicts the action [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Each action has multiple localization priors, leading to several candidate prediction results. NMS enables XRFMamba [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Example of Response Meaning Consistency (RMC) calculation. For each test action sequence, several actions with start [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: Visual comparison of XRFMamba and existing approaches. (a) presents a comparison of mAP across different tIoU [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: Example of the predicted results in Bedroom, Study room, and Dining room. In all the scenes, XRFMamba demonstrates [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: RMC scores of three LLM agents across auditor groups. XRFMamba shows consistently high performance around [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: XRFMamba, when combined with LLMs, can function as an intelligent ambient sensing agent, such as a personal [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]
Figure 14
Figure 14. Figure 14: Performance of XRFMamba in a leave-one-person-out evaluation. [PITH_FULL_IMAGE:figures/full_fig_p022_14.png]
Figure 15
Figure 15. Figure 15: The XRF V2 dataset allows combinations of different devices, and we tested the most common device combinations, [PITH_FULL_IMAGE:figures/full_fig_p024_15.png]
Figure 16
Figure 16. Figure 16: Device combinations results. As more devices are added, XRFMamba’s performance improves. [PITH_FULL_IMAGE:figures/full_fig_p025_16.png]
Figure 17
Figure 17. Figure 17: We recruited 5 volunteers to conduct a zero-shot evaluation in a real apartment. [PITH_FULL_IMAGE:figures/full_fig_p025_17.png]
Figure 18
Figure 18. Figure 18: XRF V2 already includes 2D pose data, processed using OpenPose’s Body25 model [ [PITH_FULL_IMAGE:figures/full_fig_p027_18.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-Domain Multi-Person Human Activity Recognition via Near-Field Wi-Fi Sensing

    eess.SP 2025-09 conditional novelty 6.0 of 10

    With 10 samples per available activity from a new user, WiAnchor recognizes 10 near-field Wi-Fi activities at 90.4% overall, including 86.3% for the two activities never fine-tuned.

Reference graph

Works this paper leans on

94 extracted references · 66 canonical work pages · cited by 1 Pith paper

  1. [35]

    Zhendong Liu, Le Zhang, Bing Li, Yingjie Zhou, Zhenghua Chen, and Ce Zhu. 2025. WiFi CSI Based Temporal Activity Detection Via Dual Pyramid Network. In The 39th Annual AAAI Conference on Artificial Intelligence

  2. [1]

    Mohammad Arif Ul Alam. 2020. Ai-fairness towards activity recognition of older adults. In 17th EAI International Conference on Mobile and Ubiquitous Systems: Computing, Networking and Services . 108–117

  3. [2]

    Kamran Ali, Alex X Liu, Wei Wang, and Muhammad Shahzad. 2015. Keystroke recognition using wifi signals. In Proceedings of the 21st annual international conference on mobile computing and networking . 90–102

  4. [3]

    Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, and Jorge Luis Reyes-Ortiz. 2013. A public domain dataset for human activity recognition using smartphones.. In European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, Vol. 3. 3

  5. [4]

    Behrooz Azadi, Michael Haslgrübler, Georgios Sopidis, Michaela Murauer, Bernhard Anzengruber, and Alois Ferscha. 2019. Feasibility analysis of unsupervised industrial activity recognition based on a frequent micro action. In Proceedings of the 12th ACM International Conference on PErvasive Technologies Related to Assistive Environments . 368–375

  6. [5]

    Paramvir Bahl and Venkata N Padmanabhan. 2000. RADAR: An in-building RF-based user location and tracking system. In Proceedings IEEE INFOCOM 2000. Conference on computer communications. Nineteenth annual joint conference of the IEEE computer and communications societies (Cat. No. 00CH37064) , Vol. 2. Ieee, 775–784

  7. [6]

    Marius Bock, Michael Moeller, and Kristof Van Laerhoven. 2024. Temporal Action Localization for Inertial-based Human Activity Recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 4 (2024), 1–19

  8. [7]

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition . 961–970

Show all 94 references
  1. [8]

    Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2019. Openpose: Realtime multi-person 2d pose estimation using part affinity fields. IEEE transactions on pattern analysis and machine intelligence 43, 1 (2019), 172–186

  2. [9]

    Guo Chen, Yifei Huang, Jilan Xu, Baoqi Pei, Zhe Chen, Zhiqi Li, Jiahao Wang, Kunchang Li, Tong Lu, and Limin Wang. 2024. Video mamba suite: State space model as a versatile alternative for video understanding. arXiv preprint arXiv:2403.09626 (2024)

  3. [10]

    Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou. 2024. A decoder-only foundation model for time-series forecasting. In Forty-first International Conference on Machine Learning

  4. [11]

    Arne De Brabandere, Jill Emmerzaal, Annick Timmermans, Ilse Jonkers, Benedicte Vanwanseele, and Jesse Davis. 2020. A machine learning approach to estimate hip and knee joint loading using a mobile phone-embedded IMU. Frontiers in bioengineering and biotechnology 8 (2020), 320

  5. [12]

    Nathan DeVrio, Vimal Mollyn, and Chris Harrison. 2023. SmartPoser: Arm Pose Estimation with a Smartphone and Smartwatch Using UWB and IMU Data. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–11

  6. [13]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...

  7. [14]

    Jiaqi Geng, Dong Huang, and Fernando De la Torre. 2022. Densepose from wifi. arXiv preprint arXiv:2301.00250 (2022)

  8. [15]

    Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752 (2023)

  9. [16]

    Albert Gu, Karan Goel, and Christopher Ré. 2021. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396 (2021)

  10. [17]

    Daniel Halperin, Wenjun Hu, Anmol Sheth, and David Wetherall. 2011. Tool release: Gathering 802.11 n traces with channel state information. ACM SIGCOMM computer communication review 41, 1 (2011), 53–53. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol., Vol. 0, No. 0, Arti...

  11. [18]

    Jing He and Wei Yang. 2022. Imar: Multi-user continuous action recognition with wifi signals. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 6, 3 (2022), 1–27

  12. [19]

    Alexander Hoelzemann, Julia Lee Romero, Marius Bock, Kristof Van Laerhoven, and Qin Lv. 2023. Hang-time HAR: a benchmark dataset for basketball activity recognition using wrist-worn inertial sensors. Sensors 23, 13 (2023), 5879

  13. [20]

    in the wild

    Haroon Idrees, Amir R Zamir, Yu-Gang Jiang, Alex Gorban, Ivan Laptev, Rahul Sukthankar, and Mubarak Shah. 2017. The thumos challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding 155 (2017), 1–23

  14. [21]

    Wenjun Jiang, Chenglin Miao, Fenglong Ma, Shuochao Yao, Yaqing Wang, Ye Yuan, Hongfei Xue, Chen Song, Xin Ma, Dimitrios Koutsonikolas, Wenyao Xu, and Lu Su. 2018. Towards environment independent device free human activity recognition. In Proceedings of the 24th annual internat...

  15. [22]

    Wenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang, Sen Lin, Chong Tian, Srinivasan Murali, Haochen Hu, Zhi Sun, and Lu Su

  16. [23]

    Woosub Jung, Kenneth Koltermann, Noah Helm, GinaMari Blackwell, Ingrid Pretzer-Aboff, Leslie Cloud, and Gang Zhou. 2022. IMU Sensing Data-Based Kinetic Tremor Detection in Parkinson’s Disease Patients. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Syst...

  17. [24]

    Manikanta Kotaru, Kiran Joshi, Dinesh Bharadia, and Sachin Katti. 2015. Spotfi: Decimeter level localization using wifi. In Proceedings of the 2015 ACM conference on special interest group on data communication . 269–282

  18. [25]

    Hilde Kuehne, Ali Arslan, and Thomas Serre. 2014. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the IEEE conference on computer vision and pattern recognition . 780–787

  19. [26]

    Gierad Laput and Chris Harrison. 2019. Sensing fine-grained hand activity with smartwatches. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–13

  20. [27]

    Lik-Hang Lee and Pan Hui. 2018. Interaction methods for smart glasses: A survey. IEEE Access 6 (2018), 28712–28732

  21. [28]

    Hong Li, Wei Yang, Jianxin Wang, Yang Xu, and Liusheng Huang. 2016. WiFinger: Talk to your smart devices with finger-grained gesture. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing . 250–261

  22. [29]

    Shengjie Li, Zhaopeng Liu, Qin Lv, Yanyan Zou, Yue Zhang, and Daqing Zhang. 2024. WiLife: Long-term Daily Status Monitoring and Habit Mining of the Elderly Leveraging Ubiquitous Wi-Fi Signals. ACM Transactions on Computing for Healthcare (2024)

  23. [30]

    Xiang Li, Shengjie Li, Daqing Zhang, Jie Xiong, Yasha Wang, and Hong Mei. 2016. Dynamic-MUSIC: Accurate device-free indoor localization. In Proceedings of the 2016 ACM international joint conference on pervasive and ubiquitous computing . 196–207

  24. [31]

    Jensen, and Bin Yang

    Zhe Li, Xiangfei Qiu, Peng Chen, Yihang Wang, Hanyin Cheng, Yang Shu, Jilin Hu, Chenjuan Guo, Aoying Zhou, Qingsong Wen, Christian S. Jensen, and Bin Yang. 2024. Foundts: Comprehensive and unified benchmarking of foundation models for time series forecasting. arXiv preprint ar...

  25. [32]

    Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie

    Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2117–2125

  26. [33]

    Girshick, Kaiming He, and Piotr Dollar

    Tsung-Yi Lin, Piotr Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollar. 2017. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV) . 2980–2988

  27. [34]

    Yi Liu, Limin Wang, Yali Wang, Xiao Ma, and Yu Qiao. 2022. Fineaction: A fine-grained video dataset for temporal action localization. IEEE transactions on image processing 31 (2022), 6937–6950

  28. [36]

    Yongsen Ma, Gang Zhou, Shuangquan Wang, Hongyang Zhao, and Woosub Jung. 2018. SignFi: Sign language recognition using WiFi. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 1 (2018), 1–21

  29. [37]

    Huina Meng, Xilei Wu, Xin Wang, Yuhan Fan, Jingang Shi, Han Ding, and Fei Wang. 2022. Mask wearing status estimation with smartwatches. arXiv preprint arXiv:2205.06113 (2022)

  30. [38]

    Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison, and Karan Ahuja. 2023. Imuposer: Full-body pose estimation using imus in phones, watches, and earbuds. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–12

  31. [39]

    Jay Prakash, Zhijian Yang, Yu-Lin Wei, Haitham Hassanieh, and Romit Roy Choudhury. 2020. EarSense: earphones as a teeth activity sensor. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking . 1–13

  32. [40]

    Qifan Pu, Sidhant Gupta, Shyamnath Gollakota, and Shwetak Patel. 2013. Whole-home gesture recognition using wireless signals. In Proceedings of the 19th annual international conference on Mobile computing & networking . 27–38

  33. [41]

    Philipp M Scholl, Matthias Wille, and Kristof Van Laerhoven. 2015. Wearables in the wet lab: a laboratory system for capturing and guiding experiments. In Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing . 589–599

  34. [42]

    Meng Shang, Lenore Dedeyne, Jolan Dupont, Laura Vercauteren, Nadjia Amini, Laurence Lapauw, Evelien Gielen, Sabine Verschueren, Carolina Varon, Walter De Raedt, and Bart Vanrumste. 2024. Otago exercises monitoring for older adults by a single imu and hierarchical machine learn...

  35. [43]

    Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Li, and Dacheng Tao. 2023. Tridet: Temporal action detection with relative boundary modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18857–18866

  36. [44]

    Xiaoming Shi, Shiyu Wang, Yuqi Nie, Dianqi Li, Zhou Ye, Qingsong Wen, and Ming Jin. 2024. Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts. arXiv preprint arXiv:2409.16040 (2024)

  37. [45]

    Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016. Temporal action localization in untrimmed videos via multi-stage cnns. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1049–1058

  38. [46]

    Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. 2022. Simplified state space layers for sequence modeling. arXiv preprint arXiv:2208.04933 (2022)

  39. [47]

    Georgios Sopidis, Michael Haslgrübler, Behrooze Azadi, Bernhard Anzengruber-Tánase, Abdelrahman Ahmad, Alois Ferscha, and Martin Baresch. 2022. Micro-activity recognition in industrial assembly process with IMU data and deep learning. In Proceedings of the 15th International C...

  40. [48]

    Sebastian Stein and Stephen J McKenna. 2013. Combining embedded accelerometers with computer vision for recognizing food preparation activities. In Proceedings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing . 729–738

  41. [49]

    Ke Sun, Chunyu Xia, Xinyu Zhang, Hao Chen, and Charlie Jianzhong Zhang. 2024. Multimodal Daily-Life Logging in Free-living Environment Using Non-Visual Egocentric Sensors on a Smartphone. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 ...

  42. [50]

    Sheng Tan and Jie Yang. 2016. WiFinger: Leveraging commodity WiFi for fine-grained finger gesture recognition. In Proceedings of the 17th ACM international symposium on mobile ad hoc networking and computing . 201–210

  43. [51]

    Tuan N Tang, Kwonyoung Kim, and Kwanghoon Sohn. 2023. Temporalmaxer: Maximize temporal context with only max pooling for temporal action localization. arXiv preprint arXiv:2303.09055 (2023)

  44. [52]

    Jiacheng Tian, Pan Zhou, Fangmin Sun, Tao Wang, and Hexiang Zhang. 2021. Wearable IMU-based gym exercise recognition using data fusion methods. In The Fifth International Conference on Biological Information and Biomedical Engineering . 1–7

  45. [53]

    Yonglong Tian, Guang-He Lee, Hao He, Chen-Yu Hsu, and Dina Katabi. 2018. RF-based fall monitoring using convolutional neural networks. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2, 3 (2018), 1–24

  46. [54]

    Dhruv Verma, Sejal Bhalla, Dhruv Sahnan, Jainendra Shukla, and Aman Parnami. 2021. Expressear: Sensing fine-grained facial expressions with earables. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 3 (2021), 1–28

  47. [55]

    Florian Wahl, Martin Freund, and Oliver Amft. 2015. Using smart eyeglasses as a wearable game controller. In Adjunct Proceedings of the 2015 ACM International Joint Conference on Pervasive and Ubiquitous Computing and Proceedings of the 2015 ACM International Symposium on Wear...

  48. [56]

    Fei Wang, Jianwei Feng, Yinliang Zhao, Xiaobin Zhang, Shiyuan Zhang, and Jinsong Han. 2019. Joint Activity Recognition and Indoor Localization With WiFi Fingerprints. IEEE Access 7 (2019), 80058–80068

  49. [57]

    Fei Wang, Yiao Gao, Bo Lan, Han Ding, Jingang Shi, and Jinsong Han. 2023. U-Shape Networks Are Unified Backbones for Human Action Understanding From Wi-Fi Signals. IEEE Internet of Things Journal (2023)

  50. [58]

    Fei Wang, Yizhe Lv, Mengdie Zhu, Han Ding, and Jinsong Han. 2024. XRF55: A Radio Frequency Dataset for Human Indoor Action Analysis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 1 (2024), 1–34

  51. [59]

    Fei Wang, Stanislav Panev, Ziyi Dai, Jinsong Han, and Dong Huang. 2019. Can WiFi estimate person pose?arXiv preprint arXiv:1904.00277 (2019)

  52. [60]

    Fei Wang, Tingting Zhang, Xilei Wu, Pengcheng Wang, Xin Wang, Han Ding, Jingang Shi, Jinsong Han, and Dong Huang. 2025. You Can Wash Hands Better: Accurate Daily Handwashing Assessment with a Smartwatch. IEEE Transactions on Mobile Computing (2025)

  53. [61]

    Fei Wang, Sanping Zhou, Stanislav Panev, Jinsong Han, and Dong Huang. 2019. Person-in-WiFi: Fine-grained person perception using WiFi. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5452–5461

  54. [62]

    Hao Wang, Daqing Zhang, Junyi Ma, Yasha Wang, Yuxiang Wang, Dan Wu, Tao Gu, and Bing Xie. 2016. Human respiration detection with commodity WiFi devices: Do user location and body orientation matter?. InProceedings of the 2016 ACM international joint conference on pervasive and...

  55. [63]

    Wei Wang, Alex X Liu, Muhammad Shahzad, Kang Ling, and Sanglu Lu. 2015. Understanding and modeling of wifi signal based human activity recognition. In Proceedings of the 21st annual international conference on mobile computing and networking . 65–76

  56. [64]

    Wei Wang, Alex X Liu, Muhammad Shahzad, Kang Ling, and Sanglu Lu. 2017. Device-free human activity recognition using commercial WiFi devices. IEEE Journal on Selected Areas in Communications 35, 5 (2017), 1118–1131

  57. [65]

    Xin Wang, Xilei Wu, Huina Meng, Yuhan Fan, Jingang Shi, Han Ding, and Fei Wang. 2022. Social distancing alert with smartwatches. arXiv preprint arXiv:2205.06110 (2022)

  58. [66]

    Xuyu Wang, Chao Yang, and Shiwen Mao. 2017. TensorBeat: Tensor decomposition for monitoring multiperson breathing beats with commodity WiFi. ACM Transactions on Intelligent Systems and Technology (TIST) 9, 1 (2017), 1–27

  59. [67]

    Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, Sen Xing, Guo Chen, Junting Pan, Jiashuo Yu, Yali Wang, Limin Wang, and Yu Qiao. 2022. Internvideo: General video foundation models via generative and discriminativ...

  60. [68]

    Yan Wang, Jian Liu, Yingying Chen, Marco Gruteser, Jie Yang, and Hongbo Liu. 2014. E-eyes: Device-free location-oriented activity identification using fine-grained WiFi signatures. In Proceedings of the 20th annual international conference on Mobile computing and networking. 617–628

  61. [69]

    Yichao Wang, Yili Ren, Yingying Chen, and Jie Yang. 2022. Wi-mesh: A wifi vision-based approach for 3d human mesh construction. In Proceedings of the 20th ACM Conference on Embedded Networked Sensor Systems . 362–376

  62. [70]

    Yichao Wang, Yili Ren, Yingying Chen, and Jie Yang. 2022. A wifi vision-based 3D human mesh reconstruction. In Proceedings of the 28th Annual International Conference on Mobile Computing and Networking . 814–816

  63. [71]

    Yuxi Wang, Kaishun Wu, and Lionel M Ni. 2016. Wifall: Device-free fall detection by wireless networks. IEEE Transactions on Mobile Computing 16, 2 (2016), 581–594

  64. [72]

    Johann P Wolff, Florian Grützmacher, Arne Wellnitz, and Christian Haubelt. 2018. Activity recognition using head worn inertial sensors. In Proceedings of the 5th international Workshop on Sensor-based Activity Recognition and Interaction . 1–7

  65. [73]

    Chenshu Wu, Zheng Yang, Yunhao Liu, and Wei Xi. 2012. WILL: Wireless indoor localization without site survey. IEEE Transactions on Parallel and Distributed systems 24, 4 (2012), 839–848

  66. [74]

    Kaishun Wu, Haoyu Tan, Hoilun Ngan, Yunhuai Liu, and Lionel M Ni. 2011. Chip error pattern analysis in IEEE 802.15. 4. IEEE Transactions on Mobile Computing 11, 4 (2011), 543–552

  67. [75]

    Kaishun Wu, Jiang Xiao, Youwen Yi, Min Gao, and Lionel M Ni. 2012. FILA: Fine-grained indoor localization. In 2012 Proceedings IEEE INFOCOM. IEEE, 2210–2218

  68. [76]

    Wei Xi, Jizhong Zhao, Xiang-Yang Li, Kun Zhao, Shaojie Tang, Xue Liu, and Zhiping Jiang. 2014. Electronic frog eye: Counting crowd using WiFi. In IEEE INFOCOM 2014-IEEE Conference on Computer Communications . IEEE, 361–369

  69. [77]

    Rui Xiao, Jianwei Liu, Jinsong Han, and Kui Ren. 2021. OneFi: One-Shot Recognition for Unseen Gesture via COTS WiFi. In Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems . 206–219

  70. [78]

    Wentao Xie, Qian Zhang, and Jin Zhang. 2021. Acoustic-based upper facial action recognition for smart eyewear. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 2 (2021), 1–28

  71. [79]

    Yaxiong Xie, Zhenjiang Li, and Mo Li. 2015. Precise power delay profiling with commodity WiFi. In Proceedings of the 21st Annual international conference on Mobile Computing and Networking . 53–64

  72. [80]

    Leiyang Xu, Xiaolong Zheng, Xinrun Du, Liang Liu, and Huadong Ma. 2024. WiCamera: Vortex Electromagnetic Wave-Based WiFi Imaging. IEEE Transactions on Mobile Computing (2024)

  73. [81]

    Kangwei Yan, Fei Wang, Bo Qian, Han Ding, Jinsong Han, and Xing Wei. 2024. Person-in-WiFi 3D: End-to-End Multi-Person 3D Pose Estimation with Wi-Fi. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 969–978

  74. [82]

    Jianfei Yang, He Huang, Yunjiao Zhou, Xinyan Chen, Yuecong Xu, Shenghai Yuan, Han Zou, Chris Xiaoxuan Lu, and Lihua Xie. 2024. Mm-fi: Multi-modal non-intrusive 4d human dataset for versatile wireless sensing. Advances in Neural Information Processing Systems 36 (2024)

  75. [83]

    Zheng Yang, Zimu Zhou, and Yunhao Liu. 2013. From RSSI to CSI: Indoor localization via channel response. ACM Computing Surveys (CSUR) 46, 2 (2013), 1–32

  76. [84]

    Bohan Yu, Yuxiang Wang, Kai Niu, Youwei Zeng, Tao Gu, Leye Wang, Cuntai Guan, and Daqing Zhang. 2021. WiFi-sleep: Sleep stage monitoring using commodity Wi-Fi devices. IEEE internet of things journal 8, 18 (2021), 13900–13913

  77. [85]

    Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022. Actionformer: Localizing moments of actions with transformers. In European Conference on Computer Vision. Springer, 492–510

  78. [86]

    Jie Zhang, Zhanyong Tang, Meng Li, Dingyi Fang, Petteri Nurmi, and Zheng Wang. 2018. CrossSense: Towards cross-site and large-scale WiFi sensing. In Proceedings of the 24th annual international conference on mobile computing and networking . 305–320

  79. [87]

    Guangrong Zhao, Yiran Shen, Feng Li, Lei Liu, Lizhen Cui, and Hongkai Wen. 2024. Ui-Ear: On-face Gesture Recognition Through On-ear Vibration Sensing. IEEE Transactions on Mobile Computing (2024)

  80. [88]

    Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. 2019. Hacs: Human action clips and segments dataset for recognition and temporal localization. In Proceedings of the IEEE International Conference on Computer Vision . 8668–8678

  81. [89]

    SHENGDONG ZHAO, FELICIA TAN, and KATHERINE FENNEDY. 2023. Heads-Up Computing. Commun. ACM 66, 9 (2023)

  82. [90]

    Xiaolong Zheng, Jiliang Wang, Longfei Shangguan, Zimu Zhou, and Yunhao Liu. 2016. Smokey: Ubiquitous smoking detection with commercial WiFi infrastructures. In IEEE INFOCOM 2016-The 35th Annual IEEE International Conference on Computer Communications . IEEE, 1–9

  83. [91]

    Yue Zheng, Yi Zhang, Kun Qian, Guidong Zhang, Yunhao Liu, Chenshu Wu, and Zheng Yang. 2019. Zero-effort cross-domain gesture recognition with Wi-Fi. In Proceedings of the 17th annual international conference on mobile systems, applications, and services . 313–325

  84. [92]

    Luowei Zhou, Chenliang Xu, and Jason Corso. 2018. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 32

  85. [93]

    Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv preprint arXiv:2401.09417 (2024). Proc. ACM Interact. Mob. Wearable Ubiquitous Technol....

  86. [2020]

    In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking

    Towards 3D human pose construction using WiFi. In Proceedings of the 26th Annual International Conference on Mobile Computing and Networking. 1–14

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.