REVIEW 3 major objections 5 minor 55 references
TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading
T0 review · 3 major / 5 minor · reviewed 2026-07-10 · grok-4.5
Pith's one-line read Event-based lip reading improves when each lip location keeps its own motion trajectory before spatial pooling and when viseme sequence supervision guides the final word aggregation.
desk verdict Solid, well-ablated SOTA bump on DVS-Lip via two concrete modules (local BiMamba then differential pooling; CTC viseme decoder + gated fusion); modest subfield progress, not a paradigm shift. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Trajectory-Aware Differential Aggregation (TDA) plus Viseme-Guided Aggregation (VGA). TDA reshapes high-level features into per-location temporal sequences, runs shared BiMamba, then weights each location by its difference from the global spatial context. VGA couples a CTC viseme decoder (segment generation and modeling) with a learnable gate that interpolates mean pooling and attention driven by the CTC segment context, so articulatory structure both regularizes training and steers the final word representation.
What would settle it
Retrain and re-evaluate the identical TVTA architecture after replacing the dictionary-derived viseme targets with randomly permuted or shuffled sequences of the same length; if both overall accuracy and the learned gate's preference for the viseme-guided path drop sharply relative to the true mapping, the articulatory-supervision claim holds; if performance is essentially unchanged, the viseme path is not carrying the claimed structure.
Extended reading notes
Core claim
The authors claim that event-based isolated-word lip reading is limited by premature spatial compression and purely word-level objectives, and that two complementary temporal stages fix it: Trajectory-Aware Differential Aggregation encodes a temporal trajectory at each spatial location with a shared bidirectional Mamba before differential spatial pooling, and Viseme-Guided Aggregation uses a CTC decoder on dictionary-derived viseme sequences both as sequence supervision and as segment context that gates attention-based versus mean temporal pooling of the word encoder. With EMA teacher-student consistency under asymmetric event augmentations, the student reaches 77.49 percent overall accuracy
Load-bearing premise
The method assumes that every word label can be turned into a fixed, reliable viseme sequence by a pronunciation dictionary and a static conversion table, and that CTC training on those sequences yields segment context that is both phonetically meaningful and useful for gating word-level pooling.
Editorial extensions
If this is right
- Premature spatial pooling is a concrete bottleneck for sparse event lip streams; delaying it measurably preserves discriminable local trajectories.
- Word-level classification alone under-constrains articulatory structure; CTC on dictionary-mapped viseme sequences supplies usable intermediate supervision without frame labels.
- EMA consistency under strong event-specific augmentations (cutout, zoom, temporal mask) stabilizes training and improves student accuracy over the teacher.
- Selective state-space models outperform GRU and LSTM for the pre-aggregation local temporal stage inside TDA.
- The same two-stage temporal design is intended to extend to continuous or sentence-level event-based visual speech recognition.
Reading between the lines
- If the fixed viseme conversion table is language- or dialect-specific, VGA gains may shrink on multilingual or heavily accented speakers, pointing toward learned or multi-lingual viseme inventories.
- Differential pre-aggregation of local trajectories is likely transferable to other sparse event tasks (gesture, gait) where fine motion paths matter more than global appearance.
- Because qualitative CTC outputs already show ordered viseme progressions, measuring alignment against forced phonetic alignments would test how much true articulatory structure is recovered.
- Pairing TDA-style pre-aggregation with continuous hybrid CTC/attention decoders is a direct next test for closing the remaining gap on confusing word pairs (Acc1).
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes TVTA, a temporally enhanced architecture for isolated-word event-based lip reading on DVS-Lip. It argues that prior pipelines compress spatial responses too early and optimize only at the word-classification level. The authors introduce Trajectory-Aware Differential Aggregation (TDA), which runs bidirectional Mamba temporal modeling at each spatial location before differential spatial pooling (Eqs. 1–6), and Viseme-Guided Aggregation (VGA), which couples a CTC decoder on dictionary-derived viseme sequences with a gated attention path that uses the CTC segment context to refine word-level temporal pooling (Eqs. 7–14). An EMA teacher–student scheme with asymmetric event augmentations supplies consistency regularization. On the official DVS-Lip split the student reaches 77.49 % overall accuracy (Acc1 67.23 %, Acc2 87.79 %), modestly above HFR-Lip and STCNet, with ablations attributing gains to DropPath, TDA, VGA, and the teacher–student term.
Significance. If the reported gains hold, the work supplies a concrete and reusable design pattern for event-based visual speech: model local trajectories before spatial aggregation, and inject an intermediate articulatory sequence objective (CTC on visemes) that also guides final pooling. The contribution is incremental rather than foundational—the absolute lift over the strongest published baselines is under one point—but the ablations (Tables 2–7), training-dynamics plots, and qualitative CTC decoding (Fig. 6) make the design choices inspectable and transferable. The paper is therefore of clear interest to the event-vision and visual-speech communities, especially as a template for continuous or sentence-level extensions the authors themselves flag.
major comments (3)
- Sec. 3.3 and the claim of “viseme-aware sequence supervision”: target sequences are obtained by a fixed pronunciation dictionary plus a static conversion table (following EventLip). No sensitivity analysis is provided for alternative mappings, language-specific tables, or coarser/finer viseme inventories. Because VGA’s gated aggregation path is conditioned on the CTC segment context, the central claim that the model learns articulatory structure (rather than a weak auxiliary regularizer) rests on this untested mapping. A controlled ablation that randomizes or coarsens the table, or that reports CTC alignment quality against a held-out phonetic reference, is needed to substantiate the articulatory interpretation.
- Table 1 / Sec. 4.3: the headline improvement over HFR-Lip is +0.43 points overall and is reported from a single training run with no standard deviation, seed average, or statistical test. Given the free parameters listed in the training schedule (λ_ctc/λ_ce, λ_kl/λ_ce, segment ratio, EMA ramp, DropPath rates) and the small absolute margin, multi-seed means and confidence intervals are required before the SOTA claim can be treated as robust.
- Evaluation scope (Sec. 4.1–4.3): all quantitative claims rest on a single isolated-word benchmark (DVS-Lip). The introduction and conclusion motivate continuous/sentence-level recognition, yet no transfer experiment, continuous-protocol proxy, or second dataset is supplied. Without at least one additional setting, it remains unclear whether TDA’s pre-aggregation temporal modeling and VGA’s segment context generalize beyond the 100-word isolated setting that produced the reported numbers.
minor comments (5)
- Author contact for Hongwei Ren is given as “ww@xxx.edu.cn”; this placeholder should be corrected before publication.
- Fig. 5 (right) plots the “pool gate” but never states whether γ is the raw scalar or the sigmoid-activated value used in Eq. (14); a one-sentence clarification would help readers interpret the rising curve.
- Table 2’s first row (baseline 74.09 %) is not explicitly defined in the text; stating the exact architecture (ResNet-18 + BiGRU mean-pool, no DropPath/TDA/VGA/EMA) would improve reproducibility.
- Related-work citations for EventLip, MSTP++, and E-Mamba are thorough, yet the discussion of continuous VSR CTC/hybrid decoders (Sec. 2.2) could briefly note why those continuous protocols are not used as secondary evaluation here.
- Notation: G_n is written both as a full tensor and as a per-token slice; a short remark that the Softmax in Eq. (5) is taken over the spatial index n would remove ambiguity.
Circularity Check
No significant circularity: standard empirical CV architecture proposal with held-out benchmark evaluation and additive ablations.
full rationale
The paper proposes three architectural/training components (TDA via BiMamba + differential pooling, VGA via CTC decoder + gated aggregation, EMA teacher-student) and evaluates them by training on the official DVS-Lip train split and reporting Acc/Acc1/Acc2 on the held-out test split (Table 1). Ablations (Tables 2-7) isolate additive gains and hyper-parameter sensitivity; the CTC targets are obtained by a fixed external pronunciation-dictionary + conversion-table mapping from word labels (Sec. 3.3, following EventLip), not fitted to the test metric. No equation equates a reported accuracy or 'prediction' to a free parameter by construction, no uniqueness theorem is imported from the authors' prior work, and self-citations appear only as ordinary related-work placement for event-point methods. The derivation chain is therefore the ordinary supervised empirical loop and is self-contained against the external benchmark.
Assumptions & free parameters
free parameters (6)
- λ_ctc / λ_ce ratio =
1.0
- λ_kl / λ_ce ratio =
1.0
- VGA segment length ℓ (or ratio 1/5) =
ℓ=5 (ratio 1/5)
- EMA decay schedule (μ_small, μ_large, e_ramp) =
0.999 / 0.9998, e_ramp=40 %
- DropPath stage probabilities =
0.05–0.20
- gate initialization γ favoring mean pool
assumptions (5)
- domain assumption Word labels can be converted to viseme sequences by a fixed pronunciation dictionary and conversion table without frame-level alignment.
- domain assumption Premature spatial average pooling erases sparse local motion trajectories that are discriminative for confusable lip words.
- domain assumption Standard CTC loss with blank tokens yields useful monotonic alignments for short isolated words.
- domain assumption EMA teacher on weakly augmented events supplies a stable soft target that improves student robustness under strong event perturbations.
- ad hoc to paper Bidirectional Mamba is an appropriate selective state-space model for short local temporal sequences at each spatial token.
invented entities (2)
-
Trajectory-Aware Differential Aggregation (TDA)
-
Viseme-Guided Aggregation (VGA)
Cite this review
Pith. "Pith review of TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading." pith.science (2026). https://pith.science/paper/7VZS673F
@misc{pith2026260708236,
author = {Pith},
title = {Pith review of: TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading},
year = {2026},
howpublished = {\url{https://pith.science/paper/7VZS673F}},
note = {Machine review of arXiv:2607.08236}
}
read the original abstract
Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and motion sensitivity of event cameras. However, existing methods typically perform spatial compression before sufficient temporal modeling, which may suppress sparse and localized motion trajectories that are crucial for distinguishing similar lip movements. Moreover, most current approaches optimize temporal representations mainly at the word-classification level, leaving the underlying articulatory structure weakly constrained. To address these limitations, we propose a temporally enhanced framework for event-based lip reading. First, we introduce Trajectory-Aware Differential Aggregation (TDA), which performs local temporal modeling at each spatial location before adaptive spatial aggregation. Second, we propose Viseme-Guided Aggregation (VGA), a unified temporal module composed of a CTC decoder and a viseme-guided gated aggregation branch, which injects viseme-aware sequence supervision and improves final temporal aggregation for word recognition. Third, we incorporate an EMA teacher--student training strategy to enhance robustness under strong event perturbations. Experiments on the DVS-Lip benchmark verify the effectiveness of the proposed design, and extensive ablation studies further validate the contributions of TDA, VGA, and teacher--student consistency. Qualitative decoding results also demonstrate that the proposed CTC-based temporal modeling learns meaningful viseme-aware structure from event streams.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Hugo Bulzomi, Marcel Schweiker, Amélie Gruel, and Jean Martinet. 2023. End-to-End Neuromorphic Lip-Reading. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 4101–4108. doi:10.1109/CVPRW59228. 2023.00431
-
[2]
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging Properties in Self-Supervised Vision Transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision. 9650–9660
work page 2021
-
[3]
Joao Carreira and Andrew Zisserman. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4724–4733. doi:10.1109/CVPR.2017.502
-
[4]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. InProceedings of the 37th International Conference on Machine Learning. 1597–1607
work page 2020
-
[5]
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing. 1724–1734. doi:10.3115/v1/D14-1179
-
[6]
Joon Son Chung and Andrew Zisserman. 2016. Lip Reading in the Wild. InAsian Conference on Computer Vision. 87–103
work page 2016
- [7]
-
[8]
Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar
Tommaso Furlanello, Zachary C. Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. 2018. Born-Again Neural Networks. InInternational Conference on Machine Learning. 1607–1616. 16 Jingrong Zheng, Hongwei Ren, and Xiangqian Wu
work page 2018
Show all 55 references
-
[9]
Derpanis, and Davide Scaramuzza
Daniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, and Davide Scaramuzza. 2019. End-to-End Learning of Representations for Asynchronous Event-Based Data. InProceedings of the IEEE/CVF International Conference on Computer Vision. 5633–5643
2019
-
[10]
Martínez-Hinarejos
Adrián Gimeno-Gómez and Carlos D. Martínez-Hinarejos. 2024. Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Eval...
2024
-
[11]
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber. 2006. Connectionist Temporal Classifica- tion: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. InProceedings of the 23rd International Conference on Machine Learning. 369–376. doi:10.1...
2006 doi
-
[12]
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Ávila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap Your Own Latent: A N...
2020
-
[13]
Albert Gu and Tri Dao. 2023. Mamba: Linear-Time Sequence Modeling with Selective State Spaces.arXiv preprint arXiv:2312.00752(2023)
2023 arXiv
-
[14]
Albert Gu, Karan Goel, and Christopher Ré. 2022. Efficiently Modeling Long Sequences with Structured State Spaces. InInternational Conference on Learning Representations
2022
-
[15]
Ankit Gupta, Albert Gu, and Jonathan Berant. 2022. Diagonal State Spaces Are as Effective as Structured State Spaces. InAdvances in Neural Information Processing Systems, Vol. 35. 22982–22994
2022
-
[16]
Hazen, Kate Saenko, Chia-Hao La, and James R
Timothy J. Hazen, Kate Saenko, Chia-Hao La, and James R. Glass. 2004. A Segment-Based Audio-Visual Speech Recognizer: Data Collection, Development, and Initial Experiments. InProceedings of the 6th International Conference on Multimodal Interfaces. 235–242
2004
-
[17]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9729–9738
2020
-
[18]
Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-Term Memory.Neural Computation9, 8 (1997), 1735–1780. doi:10.1162/neco.1997.9.8.1735
1997 doi
-
[19]
Weinberger
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger. 2016. Deep Networks with Stochastic Depth. InEuropean Conference on Computer Vision. 646–661
2016
-
[20]
Bin Jiang, Changjun Yuan, Xin Wang, Zhipeng Bao, Lin Zhu, and Binqiang Luo. 2023. Point-Voxel Absorbing Graph Representation Learning for Event Stream Based Recognition.arXiv preprint arXiv:2306.05239(2023)
2023 arXiv
-
[21]
Hao Ju, Zhedong Zheng, Xueye Zheng, Wenyue Chen, Lin Wang, Dong Wang, Huchuan Lu, and Xu Jia. 2026. Event- based Lip Reading with Triplane Fusion Network.ACM Transactions on Multimedia Computing, Communications and Applications(2026). https://api.semanticscholar.org/CorpusID:285971819
2026
-
[22]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization.International Conference on Learning Representations(2015)
2015
-
[23]
Kunchang Li, Xinhao Li, Yi Wang, Yinan He, Yali Wang, Limin Wang, and Yu Qiao. 2024. VideoMamba: State Space Model for Efficient Video Understanding. InEuropean Conference on Computer Vision
2024
-
[24]
Xiaobo Liang, Lijun Wu, Juntao Li, Yue Wang, Qi Meng, Tao Qin, Wei Chen, Min Zhang, and Tie-Yan Liu. 2021. R-Drop: Regularized Dropout for Neural Networks. InAdvances in Neural Information Processing Systems
2021
-
[25]
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, Jianbin Jiao, and Yunfan Liu
-
[26]
InAdvances in Neural Information Processing Systems, Vol
VMamba: Visual State Space Model. InAdvances in Neural Information Processing Systems, Vol. 37
-
[27]
Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu. 2021. TAM: Temporal Adaptive Module for Video Recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision. 13708–13718. doi:10.1109/ ICCV48922.2021.01345
2021
-
[28]
Yansong Peng, Yuhang Zhang, Zhe Xiong, Xiaoyan Sun, and Feng Wu. 2023. GET: Group Event Transformer for Event-Based Vision. InProceedings of the IEEE/CVF International Conference on Computer Vision. 6038–6048
2023
-
[29]
Hongwei Ren, Fei Ma, Xiaopeng Lin, Yuetong Fang, Hongxiang Huang, Yue Zhou, Yulong Huang, Haotian Fu, Ziyi Yang, Youxin Jiang, Xiangqian Wu, and Bojun Cheng. 2026. Scalable Event Cloud Network for Event-based Classification. In Forty-third International Conference on Machine L...
2026
-
[30]
Hongwei Ren, Yue Zhou, Xiaopeng Lin, Haotian Fu, Yulong Huang, Fei Ma, and Bojun Cheng. 2024. SpikePoint: An End- to-End Point-Based Spiking Neural Network for Event Cameras Action Recognition.arXiv preprint arXiv:2404.13489 (2024)
2024 arXiv
-
[31]
Hongwei Ren, Yue Zhou, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang, Yuetong Fang, Fei Ma, Hao Yu, and Bojun Cheng. 2025. Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression.IEEE Transactions on Pattern Analysis and Mac...
2025 doi
-
[32]
Yiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li, and Yandong Guo. 2022. Self-Distillation from the Last Mini-Batch for Consistency Regularization. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. 11943–11952
2022
-
[33]
Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman. 2023. Simplified State Space Layers for Sequence Modeling. InInternational Conference on Learning Representations
2023
-
[34]
Ganchao Tan, Zengyu Wan, Yang Wang, Yang Cao, and Zheng-Jun Zha. 2025. Tackling Event-Based Lip-Reading by Exploring Multigrained Spatiotemporal Clues.IEEE Transactions on Neural Networks and Learning Systems36, 5 (2025), 8279–8291. doi:10.1109/TNNLS.2024.3440495
2025 doi
-
[35]
Ganchao Tan, Yang Wang, Han Han, Yang Cao, Feng Wu, and Zheng-Jun Zha. 2022. Multi-Grained Spatio-Temporal Features Perceived Network for Event-Based Lip-Reading. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20094–20103. doi:10.1109/CVPR...
2022 doi
-
[36]
Antti Tarvainen and Harri Valpola. 2017. Mean Teachers Are Better Role Models: Weight-Averaged Consistency Targets Improve Semi-Supervised Deep Learning Results. InAdvances in Neural Information Processing Systems. 1195–1204
2017
-
[37]
Samuel Thomas, Binit Purohit, Tauhidur Rahman, Richard Rose, Mubarak Shah, and Ngan Le. 2025. VALLR: Visual ASR Language Model for Lip Reading. InProceedings of the IEEE/CVF International Conference on Computer Vision
2025
-
[38]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention Is All You Need. InAdvances in Neural Information Processing Systems, Vol. 30. 5998–6008
2017
-
[39]
Qi Wang, Yemin Zhang, Junsong Yuan, and Yebin Lu. 2019. Space-Time Event Clouds for Gesture Recognition: From RGB Cameras to Event Cameras. InIEEE Winter Conference on Applications of Computer Vision. 1826–1835. doi:10.1109/WACV.2019.00199
2019 doi
-
[40]
Wenhui Wang, Yulin Jin, Yawei Li, Ruiqi Du, Hongkun Hao, Weiqiang Wang, Ning Cheng, Xiaorong Wang, and Lei Xie
-
[41]
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization.arXiv preprint arXiv:2506.22846 (2025)
2025 arXiv
-
[42]
Yan Wang, Botao Du, Yongbin Shen, Kaixiang Wu, Guangtao Zhao, Jie Sun, and Huimin Wen. 2019. EV-Gait: Event- Based Robust Gait Recognition Using Dynamic Vision Sensors. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6351–6360. doi:10.1109/...
2019 doi
-
[43]
Zhongqian Wang, Qi She, and Aljosa Smolic. 2021. ACTION-Net: Multipath Excitation for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13214–13223
2021
-
[44]
Zhiming Yao, Wenyi Hong, Huihao Tang, Dong Wang, Nanxin Chen, Ke Zheng, Xiaochen Song, Dawei Yin, Furu Wei, and Ji Rong Wen. 2025. CR-CTC: Consistency Regularization on CTC for End-to-End Speech Recognition.arXiv preprint arXiv:2410.05101(2025)
2025 arXiv
-
[45]
Sukmin Yun, Jongjin Park, Kimin Lee, and Jinwoo Shin. 2020. Regularizing Class-Wise Predictions via Self-Knowledge Distillation. InIEEE/CVF Conference on Computer Vision and Pattern Recognition. 13876–13885
2020
-
[46]
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. 2019. Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation. InIEEE/CVF International Conference on Computer Vision. 3713–3722
2019
-
[47]
Wenhao Zhang, Jun Wang, Yong Luo, Lei Yu, Wei Yu, Zheng He, and Jialie Shen. 2025. MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading. InProceedings of the AAAI Conference on Artificial Intelligence
2025
-
[48]
Xueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue, Tianfang Xiao, Siqi Cai, Mingrui Lao, and Haizhou Li
-
[49]
In Proceedings of the 33rd ACM International Conference on Multimedia
EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph Modeling. In Proceedings of the 33rd ACM International Conference on Multimedia. doi:10.1145/3746027.3755404
-
[50]
Xueyi Zhang, Jialu Sun, Peiyin Zhu, Bowen Wang, Mingrui Lao, and Yanming Guo. 2025. Event-Based Binary Neural Networks for Efficient and Accurate Lip Reading. InProceedings of Machine Learning Research, Vol. 278. 1–10
2025
-
[51]
Xueyi Zhang, Jialu Sun, Peiyin Zhu, Tianfang Xiao, Mingrui Lao, and Yanming Guo. 2025. Mamba-Based Temporal Modeling for Event-Based Lip Reading. In2025 6th International Conference on Computer Vision, Image and Deep Learning
2025
-
[52]
Yi Zhang, Xiuping Liu, Hongchen Tan, and Xin Li. 2025. Spectrum-guided Spatial Feature Enhancement Network for event-based lip-reading.Neurocomputing636 (2025), 129974. doi:10.1016/j.neucom.2025.129974
2025 doi
-
[53]
Yi Zhang, Xiuping Liu, Jun Zhou, Yang Liu, and Hongchen Tan. 2025. Semantics-Aware High-Frequency Enhancement for Event-Based Lip-Reading.Information Sciences(2025), 123026. doi:10.1016/j.ins.2025.123026
2025 doi
-
[54]
Zhilu Zhang and Mert R. Sabuncu. 2020. Self-Distillation as Instance-Specific Label Smoothing. InAdvances in Neural Information Processing Systems
2020
-
[55]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. InProceedings of the 41st International Conference on Machine Learning (Proceedings of Machin...
2024
Reviewed July 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.