Pith. sign in

REVIEW 2 major objections 72 references

Paving the Way for Point Cloud Video Representation Learning Using A PDE Model

T0 review · 2 major / 0 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read A PDE model inspired by fluid analysis, solved under contrastive guidance between temporal and spatial embeddings, acts as a lightweight plug-in to improve point cloud video representation learning.

desk verdict MotionPDE tries to add PDE regularization to point cloud videos via contrastive spatial-temporal embeddings, but the discretization step on unordered sequences is the part that still needs to be shown. read the letter →

arxiv 2606.01604 v1 pith:6H3YHLIX submitted 2026-06-01 cs.CV

classification cs.CV
keywords pointcloudvideopartialdifferentialequationcontrastivelearningrepresentationspatial-temporalcorrelationmotionmodelingplug-and-playmodule
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that spatial-temporal correlations in sequential point cloud data can be regularized by recasting the learning task as the solution of a simplified partial differential equation drawn from fluid analysis. The solving process is steered by contrastive learning that aligns embeddings extracted at different times with those extracted at different spatial locations. This construction yields MotionPDE, a module that attaches to existing backbone networks, adds negligible parameters or computation, and supplies extra supervision that benefits both supervised and self-supervised regimes. A sympathetic reader would care because conventional flow-based methods break down on the irregular, unordered structure of point clouds, while this formulation supplies an alternative regularization pathway that respects the data's native geometry.

What carries the argument

MotionPDE module that formulates spatial-temporal correlations as a solvable PDE whose solution process is guided by contrastive learning between temporal and spatial embeddings.

What would settle it

Run a controlled experiment that measures accuracy or downstream task performance of a standard backbone on point cloud video benchmarks both with and without the MotionPDE module attached; absence of consistent gains would falsify the central claim.

Watch

Extended reading notes

Core claim

By constructing a simplified PDE inspired by fluid analysis and guiding the process of solving it with a contrastive structure between temporal embeddings and spatial embeddings, the authors obtain MotionPDE, an effective plug-and-play enhancement module that regularizes spatial-temporal correlation learning in point cloud videos while adding minimal computational overhead and parameters; the same contrastive process further unlocks self-supervised capabilities on this data type.

Load-bearing premise

Spatial-temporal correlations in unordered sequential point cloud data can be effectively captured and regularized by constructing and solving a simplified PDE inspired by fluid analysis, with the contrastive structure providing meaningful guidance.

Editorial extensions

If this is right

  • Backbone models for point cloud video tasks receive improved regularization of motion patterns at negligible extra cost.
  • The contrastive guidance mechanism supports self-supervised representation learning on sequential point cloud data.
  • The approach circumvents the failure modes of flow-based techniques when the input points lack a fixed spatial ordering.
  • The module can be inserted into existing architectures without redesigning the core network.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same PDE-plus-contrastive construction might transfer to other forms of irregular sequential data such as particle trajectories or mesh sequences.
  • If the learned embeddings encode physically plausible flow, they could serve as priors for downstream physics-informed tasks.
  • A natural next measurement would be to check whether the PDE residuals correlate with observed point velocities across different motion regimes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript proposes MotionPDE, a plug-and-play module for point cloud video representation learning. It formulates spatial-temporal correlation learning as a solvable PDE inspired by fluid analysis; the PDE solution process is guided by contrastive supervision between temporal and spatial embeddings. The method is claimed to enhance existing backbones with minimal added parameters and compute while also supporting self-supervised pretraining.

Significance. If the PDE discretization on unordered point clouds is shown to be stable, permutation-equivariant, and meaningfully constrained by the contrastive term (rather than reducing to ordinary contrastive learning), the approach could supply a principled regularization mechanism for sequential point-cloud data that avoids the correspondence problems of flow-based methods.

major comments (2)
  1. [Abstract] Abstract: the central performance claim requires that a simplified fluid-inspired PDE, once discretized on unordered point-cloud sequences, yields a well-posed evolution whose solution is steered by the contrastive alignment. No discretization operator (finite-difference, graph Laplacian, or particle scheme) or proof of permutation equivariance is supplied, leaving open whether the PDE residual is actually enforced or whether the regularization collapses to standard contrastive learning.
  2. [Abstract] Abstract: the claim that contrastive guidance between temporal and spatial embeddings 'refines' the PDE solution is load-bearing for the novelty argument, yet no equation or loss term is given that couples the contrastive objective to the PDE residual; without this coupling the PDE framing adds no new constraint beyond existing contrastive methods.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the careful reading and constructive comments on our work. We address each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central performance claim requires that a simplified fluid-inspired PDE, once discretized on unordered point-cloud sequences, yields a well-posed evolution whose solution is steered by the contrastive alignment. No discretization operator (finite-difference, graph Laplacian, or particle scheme) or proof of permutation equivariance is supplied, leaving open whether the PDE residual is actually enforced or whether the regularization collapses to standard contrastive learning.

    Authors: We agree that the manuscript does not supply an explicit discretization operator or a proof of permutation equivariance. The current version therefore leaves open whether the PDE residual is enforced beyond standard contrastive learning. We will revise the paper to add a description of the particle scheme used for discretization on unordered point clouds together with a proof of permutation equivariance and a clarification of how the residual is enforced. revision: yes

  2. Referee: [Abstract] Abstract: the claim that contrastive guidance between temporal and spatial embeddings 'refines' the PDE solution is load-bearing for the novelty argument, yet no equation or loss term is given that couples the contrastive objective to the PDE residual; without this coupling the PDE framing adds no new constraint beyond existing contrastive methods.

    Authors: We acknowledge that no explicit equation or loss term coupling the contrastive objective to the PDE residual appears in the manuscript. Without such a term the PDE framing may not add a new constraint. In the revision we will introduce the specific loss term that couples the contrastive supervision to the PDE residual and explain how this coupling supplies an additional constraint. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: PDE construction and contrastive guidance presented as independent steps

full rationale

The provided abstract and reader's summary describe the core approach as constructing a simplified fluid-inspired PDE and then using contrastive learning between temporal and spatial embeddings to guide its solution. No equations, definitions, or self-citations are quoted that reduce the claimed performance gain or regularization effect to a quantity fitted from the same data by construction, nor is any load-bearing premise justified solely by prior work from the same authors. The derivation chain is therefore self-contained against external benchmarks, with the PDE framing and contrastive supervision introduced as distinct contributions.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on the domain assumption that point cloud video correlations admit a useful simplified PDE description; no free parameters, invented entities, or additional axioms are stated in the abstract.

assumptions (1)
  • domain assumption Spatial-temporal correlations in unordered point cloud video data can be modeled by a simplified PDE inspired by fluid analysis.
    The paper states it constructs a simplified PDE based on this inspiration to regularize learning.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Paving the Way for Point Cloud Video Representation Learning Using A PDE Model." pith.science (2026). https://pith.science/paper/6H3YHLIX

@misc{pith2026260601604,
  author       = {Pith},
  title        = {Pith review of: Paving the Way for Point Cloud Video Representation Learning Using A PDE Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6H3YHLIX}},
  note         = {Machine review of arXiv:2606.01604}
}
read the original abstract

Investigating spatial-temporal correlations, specifically how spatial points vary over time, is crucial for understanding point cloud videos. Traditional methods, particularly flow-based techniques, struggle with these correlations due to the unordered spatial arrangement of sequential point cloud data. To address this challenge, we propose a novel approach that regularizes spatial-temporal correlation learning by formulating the problem as a solvable Partial Differential Equation (PDE). While PDEs have long been effective in the physical domain, their application to novel sequential data like point cloud video remains underexplored. Inspired by fluid analysis, we construct a simplified PDE, and the process of solving PDE is guided and refined by a contrastive learning structure between the temporal embeddings and the spatial embeddings. With this extra supervision, our method, named MotionPDE, serves as an effective, plug-and-play enhancement module for existing backbone models, adding minimal computational overhead and parameters. Capitalizing on the contrastive learning process, we delve deeper into the self-supervised capabilities of MotionPDE, yielding promising results that underscore its utility and adaptability in point cloud video data interpretation. The code repo with trained checkpoints will be available at https://github.com/zhh6425/motionpde.git for facilitating future research.

Figures

Figures reproduced from arXiv: 2606.01604 by the authors.

Figure 1
Figure 1. Motivation 1: Extra supervision related to spatial [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Motivation 2: Developing a PDE system for point [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Performance boots with our MotionPDE. in vision research, finding applications in tasks like image processing [6], [7], [8], [9], point cloud compression [28], and video prediction [10], [11]. 2.1.2 Solving Methods for PDEs While the PDE-solving problems have been widely ex￾plored with spectral methods [29], [30] and numerical meth￾ods [31], [32] since the last century. Recent research has explored deep learning mod… view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Overall architecture. MotionPDE applies separate pooling operations along the spatial and temporal dimensions to [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Data example of the SHREC 2017 dataset. Left: depth [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Model efficiency with scale-up setting. From [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 8
Figure 8. Figure 8: Follows the general principles of prior works [70], [71], we simulate noisy point cloud video data by injecting con￾trolled Gaussian noise into a subset of points. Specifically, we randomly select a fixed proportion p ∈ (0, 1) of the points in each frame and replace th…
Figure 7
Figure 7. Figure 7: A performance comparison on the MSRAction-3D [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Performance comparison on UTS-MHAD dataset [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Visualization on different clip samples of ”Pickup & Throw”. Clips are sampled from the same video data from the [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]
Figure 10
Figure 10. Figure 10: Visualization of ”Bend” and ”Forward kick” action clips. The clips are sampled from the test set of the MSRAction [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Visualization of the reconstruction results. The first row is the [PITH_FULL_IMAGE:figures/full_fig_p014_11.png]
Figure 12
Figure 12. Figure 12: Confusion matrix computed on MSRAction-3D. [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 4 canonical work pages

  1. [1]

    Self-scalable tanh (stan): Multi-scale solutions for physics- informed neural networks,

    R. Gnanasambandam, B. Shen, J. Chung, X. Yue, and Z. Kong, “Self-scalable tanh (stan): Multi-scale solutions for physics- informed neural networks,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 12, pp. 15 588–15 603, 2023

  2. [2]

    ClimODE: Climate and weather forecasting with physics-informed neural ODEs,

    Y. Verma, M. Heinonen, and V . Garg, “ClimODE: Climate and weather forecasting with physics-informed neural ODEs,” in The Twelfth International Conference on Learning Representations,

  3. [3]

    Available: https://openreview.net/forum?id= xuY33XhEGR

    [Online]. Available: https://openreview.net/forum?id= xuY33XhEGR

  4. [4]

    Spectral neural operators,

    V . Fanaskov and I. V . Oseledets, “Spectral neural operators,” in Doklady Mathematics, vol. 108, no. Suppl 2. Springer, 2023, pp. S226–S232

  5. [5]

    Solving high- dimensional pdes with latent spectral models,

    H. Wu, T. Hu, H. Luo, J. Wang, and M. Long, “Solving high- dimensional pdes with latent spectral models,” inInternational Conference on Machine Learning, 2023

  6. [6]

    HT-net: Hierarchical transformer based operator learning model for multiscale PDEs,

    X. Liu, B. Xu, and L. Zhang, “HT-net: Hierarchical transformer based operator learning model for multiscale PDEs,” 2023. [Online]. Available: https://openreview.net/forum? id=UY5zS0OsK2e

  7. [7]

    Learning to diffuse: A new perspective to design pdes for visual analy- sis,

    R. Liu, G. Zhong, J. Cao, Z. Lin, S. Shan, and Z. Luo, “Learning to diffuse: A new perspective to design pdes for visual analy- sis,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, no. 12, pp. 2457–2471, 2016. IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 16

  8. [8]

    Pde- constrained optimization in medical image analysis,

    A. Mang, A. Gholami, C. Davatzikos, and G. Biros, “Pde- constrained optimization in medical image analysis,”Optimization and Engineering, vol. 19, pp. 765–812, 2018

Show all 72 references
  1. [9]

    Formulating event-based image reconstruction as a linear inverse problem with deep regu- larization using optical flow,

    Z. Zhang, A. J. Yezzi, and G. Gallego, “Formulating event-based image reconstruction as a linear inverse problem with deep regu- larization using optical flow,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 7, pp. 8372–8389, 2023

  2. [10]

    Reformulating op- tical flow to solve image-based inverse problems and quantify uncertainty,

    A. Boquet-Pujadas and J.-C. Olivo-Marin, “Reformulating op- tical flow to solve image-based inverse problems and quantify uncertainty,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 5, pp. 6125–6141, 2023

  3. [11]

    Disentangling physical dynamics from unknown factors for unsupervised video prediction,

    V . L. Guen and N. Thome, “Disentangling physical dynamics from unknown factors for unsupervised video prediction,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11 474–11 484

  4. [12]

    Disentangling stochastic pde dynamics for unsupervised video prediction,

    X. Wu, J. Lu, Z. Yan, and G. Zhang, “Disentangling stochastic pde dynamics for unsupervised video prediction,”IEEE Transactions on Neural Networks and Learning Systems, pp. 1–15, 2023

  5. [13]

    PSTNet: Point spatio-temporal convolution on point cloud sequences,

    H. Fan, X. Yu, Y. Ding, Y. Yang, and M. Kankanhalli, “PSTNet: Point spatio-temporal convolution on point cloud sequences,” in International Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/forum?id=O3bqkf Puys

  6. [14]

    Learning representations and generative models for 3d point clouds,

    P . Achlioptas, O. Diamanti, I. Mitliagkas, and L. Guibas, “Learning representations and generative models for 3d point clouds,” in International conference on machine learning. PMLR, 2018, pp. 40– 49

  7. [15]

    Foldingnet: Point cloud auto-encoder via deep grid deformation,

    Y. Yang, C. Feng, Y. Shen, and D. Tian, “Foldingnet: Point cloud auto-encoder via deep grid deformation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 206– 215

  8. [16]

    Unsupervised learning of video representations using lstms,

    N. Srivastava, E. Mansimov, and R. Salakhudinov, “Unsupervised learning of video representations using lstms,” inInternational conference on machine learning. PMLR, 2015, pp. 843–852

  9. [17]

    Optical flow guided feature: A fast and robust motion representation for video action recognition,

    S. Sun, Z. Kuang, L. Sheng, W. Ouyang, and W. Zhang, “Optical flow guided feature: A fast and robust motion representation for video action recognition,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018

  10. [18]

    Flow dynamics correction for action recognition,

    L. Wang and P . Koniusz, “Flow dynamics correction for action recognition,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024, pp. 3795– 3799

  11. [19]

    Enhanced spatial stream of two-stream network using optical flow for human action recognition,

    S. Khan, A. Hassan, F. Hussain, A. Perwaiz, F. Riaz, M. Alsabaan, and W. Abdul, “Enhanced spatial stream of two-stream network using optical flow for human action recognition,” Applied Sciences, vol. 13, no. 14, 2023. [Online]. Available: https://www.mdpi.com/2076-3417/13/14/8003

  12. [20]

    Every frame counts: Joint learning of video segmentation and optical flow,

    M. Ding, Z. Wang, B. Zhou, J. Shi, Z. Lu, and P . Luo, “Every frame counts: Joint learning of video segmentation and optical flow,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, pp. 10 713–10 720, Apr. 2020. [Online]. Available: https://ojs.aaa...

  13. [21]

    Machine learning for partial differ- ential equations,

    S. L. Brunton and J. N. Kutz, “Machine learning for partial differ- ential equations,”arXiv preprint arXiv:2303.17078, 2023

  14. [22]

    Deep hierarchical representation of point cloud videos via spatio-temporal decom- position,

    H. Fan, X. Yu, Y. Yang, and M. Kankanhalli, “Deep hierarchical representation of point cloud videos via spatio-temporal decom- position,”IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 44, no. 12, pp. 9918–9930, 2022

  15. [23]

    Point 4d transformer networks for spatio-temporal modeling in point cloud videos,

    H. Fan, Y. Yang, and M. Kankanhalli, “Point 4d transformer networks for spatio-temporal modeling in point cloud videos,” inIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR, 2021

  16. [24]

    Point spatio-temporal transformer networks for point cloud video modeling,

    ——, “Point spatio-temporal transformer networks for point cloud video modeling,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 2, pp. 2181–2192, 2023

  17. [25]

    No pain, big gain: Classify dynamic point cloud sequences with static models by fitting feature-level space-time surfaces,

    J.-X. Zhong, K. Zhou, Q. Hu, B. Wang, N. Trigoni, and A. Markham, “No pain, big gain: Classify dynamic point cloud sequences with static models by fitting feature-level space-time surfaces,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...

  18. [26]

    Geometrymotion-net: A strong two-stream baseline for 3d action recognition,

    J. Liu and D. Xu, “Geometrymotion-net: A strong two-stream baseline for 3d action recognition,”IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, no. 12, pp. 4711–4721, 2021

  19. [27]

    Geometrymotion-transformer: An end- to-end framework for 3d action recognition,

    J. Liu, J. Guo, and D. Xu, “Geometrymotion-transformer: An end- to-end framework for 3d action recognition,”IEEE Transactions on Multimedia, pp. 1–13, 2022

  20. [28]

    G. P . Tolstov,Fourier series. Courier Corporation, 2012

  21. [29]

    Pde-based progres- sive prediction framework for attribute compression of 3d point clouds,

    X. Yang, Y. Shao, S. Liu, T. H. Li, and G. Li, “Pde-based progres- sive prediction framework for attribute compression of 3d point clouds,” inProceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 9271–9281

  22. [30]

    Gottlieb and S

    D. Gottlieb and S. A. Orszag,Numerical analysis of spectral methods: theory and applications. SIAM, 1977

  23. [31]

    Fornberg,A practical guide to pseudospectral methods

    B. Fornberg,A practical guide to pseudospectral methods. Cambridge university press, 1998, no. 1

  24. [32]

    ˆSol´ın,Partial differential equations and the finite element method

    P . ˆSol´ın,Partial differential equations and the finite element method. John Wiley & Sons, 2005

  25. [33]

    Grossmann,Numerical treatment of partial differential equations

    C. Grossmann,Numerical treatment of partial differential equations. Springer, 2007

  26. [34]

    The deep ritz method: a deep learning-based numer- ical algorithm for solving variational problems,

    B. Yuet al., “The deep ritz method: a deep learning-based numer- ical algorithm for solving variational problems,”Communications in Mathematics and Statistics, vol. 6, no. 1, pp. 1–12, 2018

  27. [35]

    A deep learning framework for solv- ing forward and inverse problems of power-law fluids,

    R. Zhai, D. Yin, and G. Pang, “A deep learning framework for solv- ing forward and inverse problems of power-law fluids,”Physics of Fluids, vol. 35, no. 9, 2023

  28. [36]

    Fourier neural oper- ator for parametric partial differential equations,

    Z. Li, N. Kovachki, K. Azizzadenesheli, B. Liu, K. Bhat- tacharya, A. Stuart, and A. Anandkumar, “Fourier neural oper- ator for parametric partial differential equations,”arXiv preprint arXiv:2010.08895, 2020

  29. [37]

    Factorized fourier neural operators,

    A. Tran, A. Mathews, L. Xie, and C. S. Ong, “Factorized fourier neural operators,” inThe Eleventh International Conference on Learning Representations, 2023. [Online]. Available: https: //openreview.net/forum?id=tmIiMPl4IPa

  30. [38]

    Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net,

    W. Luo, B. Yang, and R. Urtasun, “Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net,” inProceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2018, pp. 3569–3577

  31. [39]

    4d spatio-temporal convnets: Minkowski convolutional neural networks,

    C. Choy, J. Gwak, and S. Savarese, “4d spatio-temporal convnets: Minkowski convolutional neural networks,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3075–3084

  32. [40]

    3dv: 3d dynamic voxel for action recognition in depth video,

    Y. Wang, Y. Xiao, F. Xiong, W. Jiang, Z. Cao, J. T. Zhou, and J. Yuan, “3dv: 3d dynamic voxel for action recognition in depth video,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 511–520

  33. [41]

    Meteornet: Deep learning on dynamic 3d point cloud sequences,

    X. Liu, M. Yan, and J. Bohg, “Meteornet: Deep learning on dynamic 3d point cloud sequences,” inICCV, 2019

  34. [42]

    An efficient pointlstm for point clouds based gesture recognition,

    Y. Min, Y. Zhang, X. Chai, and X. Chen, “An efficient pointlstm for point clouds based gesture recognition,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020

  35. [43]

    Pointnet++: Deep hierar- chical feature learning on point sets in a metric space,

    C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierar- chical feature learning on point sets in a metric space,”Advances in neural information processing systems, vol. 30, 2017

  36. [44]

    Pref: Predictability regularized neural mo- tion fields,

    L. Song, X. Gong, B. Planche, M. Zheng, D. Doermann, J. Yuan, T. Chen, and Z. Wu, “Pref: Predictability regularized neural mo- tion fields,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 664–681

  37. [45]

    Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems,

    T. Chen and H. Chen, “Universal approximation to nonlinear operators by neural networks with arbitrary activation functions and its application to dynamical systems,”IEEE transactions on neural networks, vol. 6, no. 4, pp. 911–917, 1995

  38. [46]

    Physics-informed machine learning,

    G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P . Perdikaris, S. Wang, and L. Yang, “Physics-informed machine learning,”Nature Reviews Physics, vol. 3, no. 6, pp. 422–440, 2021

  39. [47]

    Learning nonlinear operators via deeponet based on the universal approx- imation theorem of operators,

    L. Lu, P . Jin, G. Pang, Z. Zhang, and G. E. Karniadakis, “Learning nonlinear operators via deeponet based on the universal approx- imation theorem of operators,”Nature machine intelligence, vol. 3, no. 3, pp. 218–229, 2021

  40. [48]

    Fourier neural operator for parametric partial differential equations,

    Z. Li, N. B. Kovachki, K. Azizzadenesheli, B. liu, K. Bhattacharya, A. Stuart, and A. Anandkumar, “Fourier neural operator for parametric partial differential equations,” inInternational Conference on Learning Representations, 2021. [Online]. Available: https://openreview.net/...

  41. [49]

    Representation learning with contrastive predictive coding,

    A. v. d. Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,”arXiv preprint arXiv:1807.03748, 2018

  42. [50]

    A simple framework for contrastive learning of visual representations,

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning. PmLR, 2020, pp. 1597– 1607

  43. [51]

    Pointmapnet: Point cloud feature map network for 3d human action IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 17 recognition,

    X. Li, Q. Huang, Y. Zhang, T. Yang, and Z. Wang, “Pointmapnet: Point cloud feature map network for 3d human action IEEE TRANSACTIONS ON PATTERN ANAL YSIS AND MACHINE INTELLIGENCE 17 recognition,”Symmetry, vol. 15, no. 2, 2023. [Online]. Available: https://www.mdpi.com/2073-899...

  44. [52]

    Real-time 3-d human action recognition based on hyperpoint sequence,

    X. Li, Q. Huang, Z. Wang, T. Yang, Z. Hou, and Z. Miao, “Real-time 3-d human action recognition based on hyperpoint sequence,” IEEE Transactions on Industrial Informatics, vol. 19, no. 8, pp. 8933– 8942, 2022

  45. [53]

    3dinaction: Under- standing human actions in 3d point clouds,

    Y. Ben-Shabat, O. Shrout, and S. Gould, “3dinaction: Under- standing human actions in 3d point clouds,”arXiv preprint arXiv:2303.06346, 2023

  46. [54]

    Hyperpointnet for point cloud sequence-based 3d human action recognition,

    X. Li, Q. Huang, T. Yang, and Q. Wu, “Hyperpointnet for point cloud sequence-based 3d human action recognition,” in2022 IEEE International Conference on Multimedia and Expo (ICME), 2022, pp. 1–6

  47. [55]

    Point primitive transformer for long-term 4d point cloud video understanding,

    H. Wen, Y. Liu, J. Huang, B. Duan, and L. Yi, “Point primitive transformer for long-term 4d point cloud video understanding,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 19–35

  48. [56]

    Point contrastive prediction with semantic clustering for self-supervised learning on point cloud videos,

    X. Sheng, Z. Shen, G. Xiao, L. Wang, Y. Guo, and H. Fan, “Point contrastive prediction with semantic clustering for self-supervised learning on point cloud videos,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 515–16 524

  49. [57]

    Mamba4d: Efficient 4d point cloud video understand- ing with disentangled spatial-temporal state space models,

    J. Liu, J. Han, L. Liu, A. I. Aviles-Rivero, C. Jiang, Z. Liu, and H. Wang, “Mamba4d: Efficient 4d point cloud video understand- ing with disentangled spatial-temporal state space models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...

  50. [58]

    Kan- hyperpointnet for point cloud sequence-based 3d human action recognition,

    Z. Chen, X. Li, Q. Huang, Q. Geng, T. Yang, and S. Han, “Kan- hyperpointnet for point cloud sequence-based 3d human action recognition,” inICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  51. [59]

    Action recognition based on a bag of 3d points,

    W. Li, Z. Zhang, and Z. Liu, “Action recognition based on a bag of 3d points,” in2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Workshops, 2010, pp. 9–14

  52. [60]

    Ntu rgb+ d: A large scale dataset for 3d human activity analysis,

    A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 1010–1019

  53. [61]

    Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,

    J. Liu, A. Shahroudy, M. Perez, G. Wang, L.-Y. Duan, and A. C. Kot, “Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding,”IEEE transactions on pattern analysis and machine intelligence, vol. 42, no. 10, pp. 2684–2701, 2019

  54. [62]

    Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor,

    C. Chen, R. Jafari, and N. Kehtarnavaz, “Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor,” in2015 IEEE International Conference on Image Processing (ICIP), 2015, pp. 168–172

  55. [63]

    Hoi4d: A 4d egocentric dataset for category- level human-object interaction,

    Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi, “Hoi4d: A 4d egocentric dataset for category- level human-object interaction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 21 ...

  56. [64]

    Shrec’17 track: 3d hand gesture recognition using a depth and skeletal dataset,

    Q. De Smedt, H. Wannous, J.-P . Vandeborre, J. Guerry, B. Le Saux, and D. Filliat, “Shrec’17 track: 3d hand gesture recognition using a depth and skeletal dataset,” in3DOR-10th Eurographics Workshop on 3D Object Retrieval, 2017, pp. 1–6

  57. [65]

    Leaf: Learning frames for 4d point cloud sequence understanding,

    Y. Liu, J. Chen, Z. Zhang, J. Huang, and L. Yi, “Leaf: Learning frames for 4d point cloud sequence understanding,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 604–613

  58. [66]

    X4d-sceneformer: Enhanced scene understanding on 4d point cloud videos through cross-modal knowledge transfer,

    L. Jing, Y. Xue, X. Yan, C. Zheng, D. Wang, R. Zhang, Z. Wang, H. Fang, B. Zhao, and Z. Li, “X4d-sceneformer: Enhanced scene understanding on 4d point cloud videos through cross-modal knowledge transfer,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 38...

  59. [67]

    Complete-to-partial 4d dis- tillation for self-supervised point cloud sequence representation learning,

    Z. Zhang, Y. Dong, Y. Liu, and L. Yi, “Complete-to-partial 4d dis- tillation for self-supervised point cloud sequence representation learning,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 17 661–17 670

  60. [68]

    Ms-tcn: Multi-stage temporal convo- lutional network for action segmentation,

    Y. A. Farha and J. Gall, “Ms-tcn: Multi-stage temporal convo- lutional network for action segmentation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 3575–3584

  61. [69]

    Ms-tcn++: Multi-stage temporal convolutional network for action segmenta- tion,

    S. Li, Y. A. Farha, Y. Liu, M.-M. Cheng, and J. Gall, “Ms-tcn++: Multi-stage temporal convolutional network for action segmenta- tion,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 6, pp. 6647–6658, 2023

  62. [70]

    Asformer: Transformer for action segmentation,

    F. Yi, H. Wen, and T. Jiang, “Asformer: Transformer for action segmentation,” inThe British Machine Vision Conference (BMVC), 2021

  63. [71]

    Pointnetlk: Robust & efficient point cloud registration using pointnet,

    Y. Aoki, H. Goforth, R. A. Srivatsan, and S. Lucey, “Pointnetlk: Robust & efficient point cloud registration using pointnet,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 7163–7172

  64. [72]

    Benchmarking and improving robustness of 3d point cloud recognition against common corruptions,

    J. Sun, Q. Zhang, B. Kailkhura, Z. Yu, C. Xiao, and Z. Mao, “Benchmarking and improving robustness of 3d point cloud recognition against common corruptions,” 2023. [Online]. Available: https://openreview.net/forum?id=wshUUnnDjc Zhuoxu Huangis a senior Ph.D. student in Computer...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.