Pith. sign in

REVIEW 3 major objections 5 minor 92 references

FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper proposes that one factorized encoder can serve as a reusable RF sensing backbone, and reports cross-domain, cross-task, and cross-modality results supporting that claim.

desk verdict Solid factorized architecture with broad evaluation, but the headline Widar3.0 gain is confounded by richer input features; the authors need a matched-input rerun. read the letter →

arxiv 2608.08439 v1 pith:4T6WJDFM submitted 2026-08-09 cs.LG cs.AI

classification cs.LGcs.AI
keywords RFsensingdomaingeneralizationcross-modalWiFiCSImmWaveradarRFIDset-basedspatialencodinghumanactivityrecognition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that a single factorized encoder, FSTC-Encoder, can serve as a reusable backbone for radio-frequency sensing across WiFi, millimeter-wave radar, and RFID, and across tasks from gesture recognition to breathing detection, even when users, environments, devices, locations, and orientations change. It argues that the obstacle to such generality is not model capacity but the coupling of behavior-relevant correlations with acquisition-specific feature geometry and spatial layout. The proposed solution standardizes the interfaces between feature, spatial, and temporal correlation stages rather than standardizing the raw inputs themselves, so only the feature configuration and task head change while the spatial-temporal backbone stays fixed. If true, this would let one sensing model be deployed in new settings without per-dataset redesign.

What carries the argument

The load-bearing mechanism is the factorization itself, enforced by three interfaces. First, structure-aware feature encoding separates branch-specific operators for each internal dependency structure (projection plus depthwise temporal convolution for low-dimensional descriptors; axis-wise convolution and masked pooling for ordered axes; paired embedding plus aggregation for coupled components) and then fuses them with sample-level gating and a joint projection. Second, set-based spatial encoding treats all spatial observations as exchangeable: each unit goes through the same residual encoder, and aggregation combines three order-invariant statistics (mean, max, standard deviation) with a set attention whose query is the spatial mean and whose values are the per-unit encodings, plus an optional receiver-quality bias. Third, hierarchical temporal encoding computes per-frame differences to expose motion transitions, applies multi-kernel depthwise convolutions, splits the sequence into overlapping windows modeled by a shared Local RoPE Transformer, and then models the window sequence with a Global RoPE Transformer using physical timestamps as scalar rotary coordinates. The output is a sequence of information-rich tokens that any lightweight task head can pool.

What would settle it

In a sensing setup where the same physical motion produces identical per-sensor signal statistics at two sensor locations but only one location's identity indicates the correct class, remove sensor identity from the input and compare accuracy with a model that receives sensor position; if the position-aware model wins, the set-based encoder's exchangeability assumption is the bottleneck. Also, permuting sensor order at test time should leave FSTC-Encoder's predictions unchanged, and any deviation would falsify the claimed permutation invariance.

Watch

Extended reading notes

Core claim

FSTC-Encoder claims that human behavior induces recurring correlations in RF signals—feature, spatial, and temporal—that can be learned separately from the measuring configuration. It decomposes representation learning into three stages: structure-aware feature encoding maps each signal family (low-dimensional descriptors, ordered axes like subcarriers or Doppler bins, pair-coupled components like I/Q or sine/cosine) into a common temporal-spatial tensor; set-based spatial encoding treats receivers, links, channels, or tags as an unordered set and aggregates them with shared per-observation encoding, order-invariant statistics, and set attention; hierarchical temporal encoding separates motion-transition enhancement, within-window local attention with rotary position encodings, and cross-window global modeling using physical timestamps. With this factorization, the paper reports 92.15% mean accuracy under its harder multi-factor cross-domain protocols on Widar3.0, the best average cross-domain Weighted-F1 on CSI-Bench, first place on three of four CSI-Bench tasks, and a reduction of the cross-modality performance gap from 18.85% to 12.93% on XRF55 through cross-RF training.

Load-bearing premise

The spatial encoding stage assumes receivers, links, channels, and tags are interchangeable: it keeps only order-invariant statistics and set attention, so any behavior-relevant information tied to a sensor's identity or geometric position must be recoverable without that identity.

Editorial extensions

If this is right

  • One trained FSTC backbone can be reused for a new RF sensing task by swapping only the feature configuration and task head, without redesigning spatial or temporal modules.
  • Cross-domain deployment becomes more predictable: the worst cross-domain performance improves along with the average, so behavior under unseen environments and users is more reliable.
  • Cross-RF learning transfers knowledge from stronger modalities like WiFi and mmWave to weaker ones like RFID, reducing the largest per-modality gap rather than only the average.
  • Input quantity alone is not the driver: with matched feature families, a general sequence model degrades when given more feature families, whereas FSTC-Encoder improves, showing that structured modeling, not scale, is doing the work.
  • Adding a new RF modality reduces to designing one structure-aware feature branch, a smaller engineering surface than rebuilding the entire model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Testable extension: freeze the spatial-temporal backbone and attach a feature branch for a completely new RF modality, such as UWB or acoustic sensing, to see whether the demonstrated transfer across WiFi, mmWave, and RFID extends further.
  • The set-based spatial encoder implies a stronger invariance claim: predictions should be identical under arbitrary permutation of receiver order, so a direct check on existing benchmarks would reveal whether residual performance differences come from non-exchangeable information like antenna geometry.
  • The cross-window physical-time modeling suggests the architecture should handle irregularly sampled temporal inputs without resampling, because timestamps enter as rotary coordinates rather than fixed grid positions.
  • The reduction of the modality gap with cross-RF training hints that the feature-spatial-temporal interfaces may also serve as a common space for unsupervised alignment across radio types, which would matter for zero-shot deployment.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. FSTC-Encoder factorizes RF sensing representation learning into feature, spatial, and temporal correlation encoding, with a structure-aware feature encoder, a set-based spatial encoder, and a hierarchical temporal encoder. The architecture is evaluated on Widar3.0, CSI-Bench, and XRF55, reporting 92.15% mean accuracy on Widar3.0 multi-factor protocols, a 57.94% CD average Weighted-F1 on CSI-Bench HAR, first place on three of four CSI-Bench tasks, and a reduction of the XRF55 cross-modality gap from 18.85% to 12.93%. The paper claims these results demonstrate domain robustness, task generality, and modality extensibility, and it uses ablations to attribute the gains to the factorized architecture rather than to richer inputs or larger models.

Significance. The factorization idea is well motivated and the architecture is clearly specified, with a clean interface between stages and a common spatial–temporal backbone across modalities. The evaluation spans multiple datasets, tasks, and RF modalities, and the matched-input ablation in Table 5 is a genuine control that goes beyond the common practice of simply adding input families. The structural ablation shows that each stage contributes. If the central claims survive the experimental concerns below, the paper would be a useful contribution to generalizable RF sensing. However, the current evidence does not yet establish the attribution claim for the headline Widar3.0 result, and the lack of uncertainty quantification makes the small reported margins difficult to interpret.

major comments (3)
  1. [Table 1 and §Cross-Domain Generalization on Widar3.0] The headline Widar3.0 comparison is not input-matched. FSTC-Encoder is fed CSI→RSSI, DFS, and dCSI, while Wi-CBR (hard) uses CSI→Phase and DFS and PatchTST uses CSI→DFS. The reported advantage over Wi-CBR (hard) is 2.95 points (92.15 vs. 89.20). Since RSSI and dCSI are additional signal quantities, the margin could be explained by additional input information rather than by the factorized architecture. The matched-input ablation in Table 5 is conducted on CSI-Bench HAR, not on Widar3.0, and it shows only that PatchTST is harmed by adding all three families in that dataset. The conclusion's sentence that matched-input analyses attribute gains to the architecture rather than richer inputs therefore overreaches for the main dataset. Please add a matched-input Widar3.0 comparison (for example, a strong baseline trained with the same three feature families) or explicitly restrict the attribution claim to the CSI-Bench setting.
  2. [All experimental tables] All empirical tables report single point estimates. For example, Table 1 reports only a mean of 92.15 for FSTC-Encoder against 89.20 for Wi-CBR (hard), and Table 2 reports a CD Average of 57.94 against 55.66 for the next-best method. No standard deviations, fold-level results, number of runs, or significance tests are given. In the absence of these, claimed margins of a few points cannot be distinguished from noise. Please report variance across folds and/or multiple seeds, and where appropriate a paired test across the cross-domain protocols.
  3. [Eqs. (9)–(13), §Set-Based Spatial Correlation Encoding] The spatial encoder is implemented as a permutation-invariant aggregation over the set of receivers, links, channels, or tags, using shared observation encoding, order-invariant statistics, and set attention with only an optional scalar quality bias β_r. This assumes that receiver identity, antenna geometry, and link topology carry no behavior-relevant information beyond what can be recovered from order-invariant statistics and the bias. The paper provides no ablation or comparison against an order-aware or geometry-aware spatial encoder (for example, one that consumes link coordinates or receiver indices) on any of the three datasets. Because the factorization argument relies on the spatial stage discarding only configuration-specific spatial layout, this assumption should be tested or at least explicitly discussed with evidence.
minor comments (5)
  1. [§Method, hyperparameters] The method section leaves the values of L_w, S_w, δ, the dimensions D_F, D_S, D_T, D_Z, the number of layers and heads, and the training procedure (optimizer, learning rate, batch size, epochs, seeds) unspecified. The free parameters δ, L_w, S_w, and β_r are named but not assigned values. Please provide these details in an appendix or supplementary material, as the empirical claims are otherwise difficult to reproduce.
  2. [§Experimental Setup, Widar3.0] The description 'holding out each condition of the target factor in turn while allowing the remaining factors to vary' does not state how many folds are used or how conditions are partitioned. Please add this detail so the reported 92.15% mean Accuracy is interpretable.
  3. [Abstract and Table 1] The phrase 'multi-factor Widar3.0 protocols' in the abstract could be read as a single protocol in which all factors vary simultaneously, whereas the experiments report a mean over cross-location, cross-orientation, and cross-environment protocols. Consider rewording to avoid ambiguity.
  4. [§In-the-Wild Domain Shifts on CSI-Bench] The sentence 'Its weaker CDev result may partly reflect information loss when reconstructing the multi-link structure from the flattened official release' is a limitation that is not further analyzed. Please provide supporting evidence or temper the claim of a balanced cross-domain profile.
  5. [Table 4 and §Modality Extensibility] The cross-RF DML training for FSTC-Multi is not described. A brief explanation of the mutual-learning objective and how the three modalities are used in training is needed to interpret the gap reduction from 18.85% to 12.93%.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the architecture is evaluated on held-out domains with matched-input and structural controls, not derived from its own assumptions.

full rationale

The paper contains no derivation that reduces to its own inputs. The method section (Eqs. 1-22) builds a feature-spatial-temporal encoder from explicit operators, and the experiments section then evaluates the resulting model on Widar3.0, CSI-Bench, and XRF55 held-out domain splits. The 92.15% figure is an observed test-set accuracy, not a quantity implied by the architecture's definition. The matched-input ablation in Table 5 compares FSTC-Encoder and PatchTST under identical feature families and shows that PatchTST degrades when all three families are combined; the structural ablation in Table 6 removes each encoding stage and measures accuracy drops. These are controls, not circular steps. No 'uniqueness theorem' or prior-work ansatz is invoked to force the architecture; related-work comparisons cite external papers such as SCL and Wi-CBR without making the present result depend on them. The references include no load-bearing self-citation chain. The closest issue is that the headline Widar3.0 comparison in Table 1 is not input-matched (FSTC uses CSI-to-RSSI, DFS, dCSI while Wi-CBR uses CSI-to-Phase, DFS), so attributing the 2.95-point gain to the architecture is underdetermined for that dataset; however, that is an experimental-confound and evidence-quality concern, not a pattern of circularity as defined by the seven enumerated kinds. No fitted parameter is renamed as a prediction, and no central claim is equivalent by construction to an input.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new physical entities, forces, or conserved quantities. The central architecture relies on standard deep-learning hyperparameters and on domain assumptions about permutation invariance of spatial units and the recoverability of physical timestamps. The only hand-chosen constant that enters the method equations is the temporal normalization unit δ, whose value is not disclosed.

free parameters (3)
  • δ (temporal normalization unit) = not disclosed
    Used in global RoPE coordinate p_j = τ_j / δ (Eq. 20). The value affects the frequency scaling of long-range temporal attention and is presumably chosen per dataset, but the paper does not state how it is set.
  • Window length L_w and stride S_w = not disclosed
    These determine the segmentation in the hierarchical temporal encoder (Eq. 16). They control the trade-off between local and global modeling, but their exact values are not reported.
  • Quality bias β_r = not disclosed
    Optional signed additive quality bias in set attention (Eq. 11), derived from receiver-level quality metadata. The derivation or learned values are not specified.
assumptions (3)
  • domain assumption Spatial observations (receivers, links, tags) are exchangeable and can be aggregated as an unordered set.
    Entered in 'Set-Based Spatial Correlation Encoding' (Eqs. 9-13). If receiver identity or geometric ordering matters beyond a quality bias, the permutation-invariant statistics and attention will discard that information.
  • domain assumption Physical timestamps are available or recoverable from frame indices and the nominal sampling interval.
    Entered in 'Cross-Window Physical-Time Modeling' after Eq. 20. If sampling is irregular or timestamps are missing, the global RoPE coordinate p_j = τ_j / δ will be inaccurate.
  • domain assumption Source and target domains share the label space but differ in class-conditional distributions.
    Standard domain generalization assumption stated in Eq. 3. If the label space differs across configurations, the task head must change accordingly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing." pith.science (2026). https://pith.science/paper/4T6WJDFM

@misc{pith2026260808439,
  author       = {Pith},
  title        = {Pith review of: FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4T6WJDFM}},
  note         = {Machine review of arXiv:2608.08439}
}
read the original abstract

Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult to reuse across devices, environments, and RF modalities. We propose FSTC-Encoder, which unifies heterogeneous RF representation learning through feature, spatial, and temporal correlation modeling. Structure-aware feature encoding accommodates different signal structures, set-based spatial encoding aggregates variable observations, and hierarchical temporal encoding jointly captures local variations and long-range dependencies. Across sensing tasks and modalities, FSTC-Encoder retains the same spatial--temporal backbone architecture while varying only the feature configuration and task head. Across Widar3.0, CSI-Bench, and XRF55, FSTC-Encoder achieves 92.15% mean Accuracy under multi-factor cross-domain protocols, ranks first on three of four additional sensing tasks, remains consistently strong across WiFi, millimeter-wave radar, and RFID, and reduces the cross-modality performance gap from 18.85% to 12.93% through cross-RF learning. These results demonstrate that FSTC-Encoder achieves high domain robustness, task generality, and modality extensibility.

Figures

Figures reproduced from arXiv: 2608.08439 by the authors.

Figure 1
Figure 1. Feature, spatial, and temporal correlations in RF [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of FSTC-Encoder. Structure-aware feature encoding maps heterogeneous inputs to a common interface, set [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 70 canonical work pages

  1. [1]

    Abdelnasser, H.; Youssef, M.; and Harras, K. A. 2015. WiGest : A Ubiquitous WiFi -Based Gesture Recognition System. In Proceedings of IEEE INFOCOM, 1472--1480

  2. [2]

    Adib, F.; Kabelac, Z.; Katabi, D.; and Miller, R. C. 2014. 3D Tracking via Body Radio Reflections. In Proceedings of the 11th USENIX Symposium on Networked Systems Design and Implementation, 317--329

  3. [3]

    Adib, F.; Mao, H.; Kabelac, Z.; Katabi, D.; and Miller, R. C. 2015. Smart Homes that Monitor Breathing and Heart Rate. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, 837--846

  4. [4]

    Bertasius, G.; Wang, H.; and Torresani, L. 2021. Is Space-Time Attention All You Need for Video Understanding? In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, 813--824

  5. [5]

    Chen, X.; and Yang, J. 2025. X-Fi : A Modality-Invariant Foundation Model for Multimodal Human Sensing. In International Conference on Learning Representations

  6. [6]

    Chen, Y.; and Huang, X. 2024. WiGNN : WiFi -Based Cross-Domain Gesture Recognition Inspired by Dynamic Topology Structure. IEEE Wireless Communications, 31(3): 249--256

  7. [7]

    N.; Hoang, N

    Dang, V. N.; Hoang, N. C.; Nguyen, Q. C.; and Le, M. T. 2025. Advancing Robust Human Activity Recognition via Informative mmWave Radar Characteristics and a Lightweight Spatio-Spectro-Temporal Network. Measurement, 256: 118056

  8. [8]

    Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations

Show all 92 references
  1. [9]

    Du, L.; Shang, S.; Zhang, L.; Li, C.; Yang, J.; and Tian, X. 2024. Multidomain Correlation-Based Multidimensional CSI Tensor Generation for Device-Free Wi-Fi Sensing. Computer Modeling in Engineering & Sciences, 138(2): 1749--1767

  2. [10]

    Fan, L.; Li, T.; Yuan, Y.; and Katabi, D. 2020. In-Home Daily-Life Captioning Using Radio Signals. In Computer Vision -- ECCV 2020, 105--123

  3. [11]

    Gao, S.; Zhang, J.; Mei, L.; Wang, S.; and Wang, X. 2026. Exploring Spatial--Temporal Representation via Star Graph for mmWave Radar-Based Human Activity Recognition. IEEE Transactions on Mobile Computing, 25(4): 5700--5715

  4. [12]

    Gu, Y.; Zhang, X.; Wang, Y.; Wang, M.; Liu, Z.; Yan, H.; Ji, Y.; Li, J.; and Dong, M. 2022. WiGRUNT : WiFi -Enabled Gesture Recognition Using Dual-Attention Network. IEEE Transactions on Human-Machine Systems, 52(4): 736--746

  5. [13]

    He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770--778

  6. [14]

    Hochreiter, S.; and Schmidhuber, J. 1997. Long Short-Term Memory. Neural Computation, 9(8): 1735--1780

  7. [15]

    Huang, S.; Li, K.; You, D.; Chen, Y.; Lin, A.; Liu, S.; Li, X.; and McCann, J. A. 2024. WiMANS : A Benchmark Dataset for WiFi -Based Multi-User Activity Sensing. In Computer Vision -- ECCV 2024, 72--91

  8. [16]

    P.; Peng, W.-H.; Ma, C.-W.; and Hwang, J.-N

    Lee, S.-P.; Kini, N. P.; Peng, W.-H.; Ma, C.-W.; and Hwang, J.-N. 2023. HuPR : A Benchmark for Human Pose Estimation Using Millimeter Wave Radar. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 5715--5724

  9. [17]

    Li, C.; Liu, M.; and Cao, Z. 2022. WiHF : Gesture and User Recognition With WiFi . IEEE Transactions on Mobile Computing, 21(2): 757--768

  10. [18]

    Li, T.; Fan, L.; Yuan, Y.; and Katabi, D. 2022. Unsupervised Learning for Human Sensing Using Radio Signals. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 3288--3297

  11. [19]

    Li, T.; Fan, L.; Zhao, M.; Liu, Y.; and Katabi, D. 2019. Making the Invisible Visible: Action Recognition Through Walls and Occlusions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 872--881

  12. [20]

    Li, X.; Chang, L.; Song, F.; Wang, J.; Chen, X.; Tang, Z.; and Wang, Z. 2021. CrossGR : Accurate and Low-Cost Cross-Target Gesture Recognition Using Wi-Fi . Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 5(1): 21:1--21:23

  13. [21]

    E.; Amihood, P.; Schwesig, C.; Olson, E.; Raja, H.; and Poupyrev, I

    Lien, J.; Gillian, N.; Karagozler, M. E.; Amihood, P.; Schwesig, C.; Olson, E.; Raja, H.; and Poupyrev, I. 2016. Soli : Ubiquitous Gesture Sensing with Millimeter Wave Radar. ACM Transactions on Graphics, 35(4): 142:1--142:19

  14. [22]

    Liu, S.; Chen, Z.; Wu, M.; Liu, C.; and Chen, L. 2024 a . WiSR : Wireless Domain Generalization Based on Style Randomization. IEEE Transactions on Mobile Computing, 23(5): 4520--4532

  15. [23]

    Liu, X.; Zhang, B.; Chen, S.; Xie, X.; Tong, X.; Gu, T.; and Li, K. 2024 b . A Wireless Signal Correlation Learning Framework for Accurate and Robust Multi-Modal Sensing. IEEE Journal on Selected Areas in Communications, 42(9): 2424--2439

  16. [24]

    Ma, Y.; Zhou, G.; Wang, S.; Zhao, H.; and Jung, W. 2018. SignFi : Sign Language Recognition Using WiFi . Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 2(1): 23:1--23:21

  17. [25]

    H.; Sinthong, P.; and Kalagnanam, J

    Nie, Y.; Nguyen, N. H.; Sinthong, P.; and Kalagnanam, J. 2023. A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers. In International Conference on Learning Representations

  18. [26]

    Pu, Q.; Gupta, S.; Gollakota, S.; and Patel, S. 2013. Whole-Home Gesture Recognition Using Wireless Signals. In Proceedings of the 19th Annual International Conference on Mobile Computing and Networking, 27--38

  19. [27]

    Ren, Q.; Wang, Y.; Liu, S.; and Lv, X. 2023. FSTNet : Learning Spatial--Temporal Correlations from Fingerprints for Indoor Positioning. Ad Hoc Networks, 149: 103244

  20. [28]

    E.; Hinton, G

    Rumelhart, D. E.; Hinton, G. E.; and Williams, R. J. 1986. Learning Representations by Back-Propagating Errors. Nature, 323(6088): 533--536

  21. [29]

    Shang, M.; and Hong, X. 2023. Recurrent ConFormer for WiFi Activity Recognition. IEEE/CAA Journal of Automatica Sinica, 10(6): 1491--1493

  22. [30]

    N.; Kaiser, L.; and Polosukhin, I

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems, volume 30

  23. [31]

    Virmani, A.; and Shahzad, M. 2017. Position and Orientation Agnostic Gesture Recognition Using WiFi . In Proceedings of the 15th Annual International Conference on Mobile Systems, Applications, and Services, 252--264

  24. [32]

    Wang, F.; Lv, Y.; Zhu, M.; Ding, H.; and Han, J. 2024. XRF55 : A Radio Frequency Dataset for Human Indoor Action Analysis. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 8(1): 21:1--21:34

  25. [33]

    Wang, F.; Zhou, S.; Panev, S.; Han, J.; and Huang, D. 2019. Person-in-WiFi : Fine-Grained Person Perception Using WiFi . In Proceedings of the IEEE/CVF International Conference on Computer Vision, 5452--5461

  26. [35]

    X.; Shahzad, M.; Ling, K.; and Lu, S

    Wang, W.; Liu, A. X.; Shahzad, M.; Ling, K.; and Lu, S. 2017. Device-Free Human Activity Recognition Using Commercial WiFi Devices. IEEE Journal on Selected Areas in Communications, 35(5): 1118--1131

  27. [37]

    Yan, H.; Zhang, X.; Huang, J.; Feng, Y.; Li, M.; Wang, A.; Ou, W.; Wang, H.; and Liu, Z. 2025. Wi-SFDAGR : WiFi -Based Cross-Domain Gesture Recognition via Source-Free Domain Adaptation. IEEE Internet of Things Journal, 12(13): 24159--24173

  28. [38]

    X.; and Xie, L

    Yang, J.; Huang, H.; Zhou, Y.; Chen, X.; Xu, Y.; Yuan, S.; Zou, H.; Lu, C. X.; and Xie, L. 2023. MM-Fi : Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing. In Advances in Neural Information Processing Systems, volume 36, 18756--18768

  29. [40]

    Yao, S.; Hu, S.; Zhao, Y.; Zhang, A.; and Abdelzaher, T. 2017. DeepSense : A Unified Deep Learning Framework for Time-Series Mobile Sensing Data Processing. In Proceedings of the 26th International Conference on World Wide Web, 351--360

  30. [41]

    Zhang, R.; Tang, S.; Yan, H.; Zhang, X.; and Guo, J. 2026. Wi-CBR : Salient-Aware Adaptive WiFi Sensing for Cross-Domain Behavior Recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 40(2): 1552--1560

  31. [42]

    Zhang, Y.; Zheng, Y.; Qian, K.; Zhang, G.; Liu, Y.; Wu, C.; and Yang, Z. 2022. Widar3.0 : Zero-Effort Cross-Domain Gesture Recognition With Wi-Fi . IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11): 8671--8688

  32. [43]

    A.; Tian, Y.; Zhao, H.; Torralba, A.; and Katabi, D

    Zhao, M.; Li, T.; Alsheikh, M. A.; Tian, Y.; Zhao, H.; Torralba, A.; and Katabi, D. 2018. Through-Wall Human Pose Estimation Using Radio Signals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 7356--7365

  33. [44]

    Zhao, M.; Liu, Y.; Raghu, A.; Zhao, H.; Li, T.; Torralba, A.; and Katabi, D. 2019. Through-Wall Human Mesh Recovery Using Radio Signals. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10113--10122

  34. [45]

    Zhu, G.; Hu, Y.; Gao, W.; Wang, W.-H.; Wang, B.; and Liu, K. J. R. 2025. CSI-Bench : A Large-Scale In-the-Wild Dataset for Multi-Task WiFi Sensing. In Advances in Neural Information Processing Systems, volume 38

  35. [46]

    Proceedings of the 19th Annual International Conference on Mobile Computing and Networking , pages =

    Pu, Qifan and Gupta, Sidhant and Gollakota, Shyamnath and Patel, Shwetak , title =. Proceedings of the 19th Annual International Conference on Mobile Computing and Networking , pages =. 2013 , doi =

  36. [47]

    , title =

    Abdelnasser, Heba and Youssef, Moustafa and Harras, Khaled A. , title =. Proceedings of IEEE INFOCOM , pages =. 2015 , doi =

  37. [48]

    and Shahzad, Muhammad and Ling, Kang and Lu, Sanglu , title =

    Wang, Wei and Liu, Alex X. and Shahzad, Muhammad and Ling, Kang and Lu, Sanglu , title =. IEEE Journal on Selected Areas in Communications , volume =. 2017 , doi =

  38. [49]

    Proceedings of the 15th Annual International Conference on Mobile Systems, Applications, and Services , pages =

    Virmani, Aditya and Shahzad, Muhammad , title =. Proceedings of the 15th Annual International Conference on Mobile Systems, Applications, and Services , pages =. 2017 , doi =

  39. [50]

    IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =

    Zhang, Yi and Zheng, Yue and Qian, Kun and Zhang, Guidong and Liu, Yunhao and Wu, Chenshu and Yang, Zheng , title =. IEEE Transactions on Pattern Analysis and Machine Intelligence , volume =. 2022 , doi =

  40. [51]

    IEEE Transactions on Mobile Computing , volume =

    Li, Chenning and Liu, Manni and Cao, Zhichao , title =. IEEE Transactions on Mobile Computing , volume =. 2022 , doi =

  41. [52]

    Proceedings of the 26th International Conference on World Wide Web , pages =

    Yao, Shuochao and Hu, Shaohan and Zhao, Yiran and Zhang, Aston and Abdelzaher, Tarek , title =. Proceedings of the 26th International Conference on World Wide Web , pages =. 2017 , doi =

  42. [53]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =

    Ma, Yongsen and Zhou, Gang and Wang, Shuangquan and Zhao, Hongyang and Jung, Woosub , title =. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =. 2018 , doi =

  43. [54]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

    Zhao, Mingmin and Li, Tianhong and Alsheikh, Mohammad Abu and Tian, Yonglong and Zhao, Hang and Torralba, Antonio and Katabi, Dina , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2018 , doi =

  44. [55]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Li, Tianhong and Fan, Lijie and Zhao, Mingmin and Liu, Yingcheng and Katabi, Dina , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =. 2019 , doi =

  45. [56]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Wang, Fei and Zhou, Sanping and Panev, Stanislav and Han, Jinsong and Huang, Dong , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =. 2019 , doi =

  46. [57]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Zhao, Mingmin and Liu, Yingcheng and Raghu, Aniruddh and Zhao, Hang and Li, Tianhong and Torralba, Antonio and Katabi, Dina , title =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =. 2019 , doi =

  47. [58]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =

    Li, Tianhong and Fan, Lijie and Yuan, Yuan and Katabi, Dina , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2022 , doi =

  48. [59]

    IEEE Transactions on Human-Machine Systems , volume =

    Gu, Yu and Zhang, Xiang and Wang, Yantong and Wang, Meng and Liu, Zhi and Yan, Huan and Ji, Yusheng and Li, Jianhua and Dong, Mianxiong , title =. IEEE Transactions on Human-Machine Systems , volume =. 2022 , doi =

  49. [60]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =

    Li, Xinyi and Chang, Liqiong and Song, Fangfang and Wang, Ju and Chen, Xiaojiang and Tang, Zhanyong and Wang, Zheng , title =. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =. 2021 , doi =

  50. [61]

    IEEE Internet of Things Journal , volume =

    Yan, Huan and Zhang, Xiang and Huang, Jinyang and Feng, Yuanhao and Li, Meng and Wang, Anzhi and Ou, Weihua and Wang, Hongbing and Liu, Zhi , title =. IEEE Internet of Things Journal , volume =. 2025 , doi =

  51. [62]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Zhang, Ruobei and Tang, Shengeng and Yan, Huan and Zhang, Xiang and Guo, Jiabao , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =

  52. [63]

    Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =

    Lee, Shih-Po and Kini, Niraj Prakash and Peng, Wen-Hsiao and Ma, Ching-Wen and Hwang, Jenq-Neng , title =. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , pages =. 2023 , doi =

  53. [64]

    IEEE Wireless Communications , volume =

    Chen, Yinan and Huang, Xiaoxia , title =. IEEE Wireless Communications , volume =. 2024 , doi =

  54. [65]

    Computer Modeling in Engineering & Sciences , volume =

    Du, Liufeng and Shang, Shaoru and Zhang, Linghua and Li, Chong and Yang, Jianing and Tian, Xiyan , title =. Computer Modeling in Engineering & Sciences , volume =. 2024 , doi =

  55. [66]

    IEEE Journal on Selected Areas in Communications , volume =

    Liu, Xiulong and Zhang, Bojun and Chen, Sheng and Xie, Xin and Tong, Xinyu and Gu, Tao and Li, Keqiu , title =. IEEE Journal on Selected Areas in Communications , volume =. 2024 , doi =

  56. [67]

    , title =

    Adib, Fadel and Kabelac, Zachary and Katabi, Dina and Miller, Robert C. , title =. Proceedings of the 11th USENIX Symposium on Networked Systems Design and Implementation , pages =

  57. [68]

    , title =

    Adib, Fadel and Mao, Hongzi and Kabelac, Zachary and Katabi, Dina and Miller, Robert C. , title =. Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems , pages =. 2015 , doi =

  58. [69]

    ACM Transactions on Graphics , volume =

    Lien, Jaime and Gillian, Nicholas and Karagozler, Mustafa Emre and Amihood, Patrick and Schwesig, Carsten and Olson, Erik and Raja, Hakim and Poupyrev, Ivan , title =. ACM Transactions on Graphics , volume =. 2016 , doi =

  59. [70]

    Computer Vision -- ECCV 2020 , pages =

    Fan, Lijie and Li, Tianhong and Yuan, Yuan and Katabi, Dina , title =. Computer Vision -- ECCV 2020 , pages =. 2020 , doi =

  60. [71]

    Advances in Neural Information Processing Systems , volume =

    Yang, Jianfei and Huang, He and Zhou, Yunjiao and Chen, Xinyan and Xu, Yuecong and Yuan, Shenghai and Zou, Han and Lu, Chris Xiaoxuan and Xie, Lihua , title =. Advances in Neural Information Processing Systems , volume =

  61. [72]

    , title =

    Huang, Shuokang and Li, Kaihan and You, Di and Chen, Yichong and Lin, Arvin and Liu, Siying and Li, Xiaohui and McCann, Julie A. , title =. Computer Vision -- ECCV 2024 , pages =. 2024 , doi =

  62. [73]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =

    Wang, Fei and Lv, Yizhe and Zhu, Mengdie and Ding, Han and Han, Jinsong , title =. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies , volume =. 2024 , doi =

  63. [74]

    Zhu, Guozhen and Hu, Yuqian and Gao, Weihang and Wang, Wei-Hsiang and Wang, Beibei and Liu, K. J. Ray , title =. Advances in Neural Information Processing Systems , volume =

  64. [75]

    Ad Hoc Networks , volume =

    Ren, Qianqian and Wang, Yan and Liu, Saining and Lv, Xingfeng , title =. Ad Hoc Networks , volume =. 2023 , doi =

  65. [76]

    IEEE Transactions on Mobile Computing , volume =

    Gao, Senhao and Zhang, Junqing and Mei, Luoyu and Wang, Shuai and Wang, Xuyu , title =. IEEE Transactions on Mobile Computing , volume =. 2026 , doi =

  66. [77]

    Measurement , volume =

    Dang, Van Ngoc and Hoang, Ngoc Chau and Nguyen, Quoc Cuong and Le, Minh Thuy , title =. Measurement , volume =. 2025 , doi =

  67. [78]

    Proceedings of the 24th Annual International Conference on Mobile Computing and Networking , pages =

    Jiang, Wenjun and Miao, Chenglin and Ma, Fenglong and Yao, Shuochao and Wang, Yaqing and Yuan, Ye and Xue, Hongfei and Song, Chen and Ma, Xin and Koutsonikolas, Dimitrios and Xu, Wenyao and Su, Lu , title =. Proceedings of the 24th Annual International Conference on Mobile Com...

  68. [79]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Li, Bing and Cui, Wei and Wang, Wei and Zhang, Le and Chen, Zhenghua and Wu, Min , title =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2021 , doi =

  69. [80]

    Advances in Neural Information Processing Systems , volume =

    Yang, Shiqi and Wang, Yaxing and Wang, Kai and Jui, Shangling and Van De Weijer, Joost , title =. Advances in Neural Information Processing Systems , volume =

  70. [81]

    IEEE Sensors Journal , volume =

    Zhang, Changsheng and Jiao, Wanguo , title =. IEEE Sensors Journal , volume =. 2023 , doi =

  71. [82]

    IEEE Transactions on Mobile Computing , volume =

    Liu, Shijia and Chen, Zhenghua and Wu, Min and Liu, Chang and Chen, Liangyin , title =. IEEE Transactions on Mobile Computing , volume =. 2024 , doi =

  72. [83]

    IEEE/CAA Journal of Automatica Sinica , volume =

    Shang, Miao and Hong, Xiaopeng , title =. IEEE/CAA Journal of Automatica Sinica , volume =. 2023 , doi =

  73. [84]

    2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing , pages =

    Dai, Miaoling and Cao, Chenhong and Liu, Tong and Su, Meijia and Li, Yufeng and Li, Jiangtao , title =. 2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing , pages =. 2023 , doi =

  74. [85]

    and Hinton, Geoffrey E

    Rumelhart, David E. and Hinton, Geoffrey E. and Williams, Ronald J. , title =. Nature , volume =. 1986 , doi =

  75. [86]

    Long Short-Term Memory , journal =

    Hochreiter, Sepp and Schmidhuber, J. Long Short-Term Memory , journal =. 1997 , doi =

  76. [87]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =

    He, Kaiming and Zhang, Xiangyu and Ren, Shaoqing and Sun, Jian , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =

  77. [88]

    and Kaiser, Lukasz and Polosukhin, Illia , title =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , title =. Advances in Neural Information Processing Systems , volume =

  78. [89]

    International Conference on Learning Representations , year =

    Dosovitskiy, Alexey and Beyer, Lucas and Kolesnikov, Alexander and Weissenborn, Dirk and Zhai, Xiaohua and Unterthiner, Thomas and Dehghani, Mostafa and Minderer, Matthias and Heigold, Georg and Gelly, Sylvain and Uszkoreit, Jakob and Houlsby, Neil , title =. International Con...

  79. [90]

    and Sinthong, Phanwadee and Kalagnanam, Jayant , title =

    Nie, Yuqi and Nguyen, Nam H. and Sinthong, Phanwadee and Kalagnanam, Jayant , title =. International Conference on Learning Representations , year =

  80. [91]

    Proceedings of the 38th International Conference on Machine Learning , series =

    Bertasius, Gedas and Wang, Heng and Torresani, Lorenzo , title =. Proceedings of the 38th International Conference on Machine Learning , series =

  81. [92]

    arXiv preprint arXiv:2504.14621 , year =

    Yang, Zhenkui and Huang, Zeyi and Wang, Ge and Ding, Han and Han, Tony Xiao and Wang, Fei , title =. arXiv preprint arXiv:2504.14621 , year =

  82. [93]

    International Conference on Learning Representations , year =

    Chen, Xinyan and Yang, Jianfei , title =. International Conference on Learning Representations , year =

  83. [94]

    arXiv preprint arXiv:2604.05584 , year =

    Weng, Pengcheng and Qian, Yanyu and Xu, Yangxin and Wang, Fei , title =. arXiv preprint arXiv:2604.05584 , year =

  84. [95]

    arXiv preprint arXiv:2604.02056 , year =

    Wang, Hao and Qian, Yanyu and Weng, Pengcheng and Xia, Zixuan and Dan, William and Xu, Yangxin and Wang, Fei , title =. arXiv preprint arXiv:2604.02056 , year =

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.