Pith. sign in

REVIEW 3 major objections 5 minor 89 references

OPEN: A Benchmark Dataset and Baseline for Older Adult Patient Engagement Recognition in Virtual Rehabilitation Learning Environments

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper introduces OPEN, the largest public dataset for AI-driven engagement recognition in older adult virtual rehabilitation patients, with baselines reaching 81% detection accuracy.

desk verdict Useful dataset, but the 81% 'engagement' accuracy is essentially on-task/off-task detection—frame it that way and it's a solid contribution. read the letter →

arxiv 2507.17959 v1 pith:RFKRSIFP submitted 2025-07-23 cs.CV

classification cs.CV
keywords engagementrecognitionvirtualrehabilitationolderadultslandmarkfeaturesaffectivecomputingcardiacmachinelearningbenchmarkprivacy-preservingdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces OPEN, a public dataset for AI-driven engagement recognition built from 35 hours and 2 minutes of virtual cardiac rehabilitation sessions attended by eleven older adults over six weeks. The authors' central claim is that OPEN is the largest publicly available engagement dataset for this population, and that it fills a demographic gap left by datasets collected from young students in single sessions. They argue that releasing facial, hand, and body landmarks instead of raw video preserves privacy while still supporting competitive models, with baselines reaching 81% accuracy on engagement detection and 70% on engagement prediction. A sympathetic reader would care because automated engagement monitoring could help clinicians support adherence in virtual rehabilitation, where dropout is a known problem.

What carries the argument

The load-bearing mechanism is the combination of the HELP annotation protocol with adaptive temporal segmentation and landmark-only feature release. HELP defines the binary ground truth as a deterministic rule over four affective and two behavioral classes, and OPEN's annotators recorded state changes to the second, so each variable-length sample is internally consistent rather than mixing engagement states. The released feature set, including OpenFace eye gaze, head pose and facial action units, EmoFAN valence-arousal, and MediaPipe facial, hand and body landmarks, is what lets models train on identifiable behavioral signals, while the adaptive segmentation is what enables both detection and the new prediction task.

What would settle it

Collect self-reported engagement or clinician-rated engagement for the same 36 sessions and compare those ratings to the HELP-derived labels; if agreement is near chance, the ground truth does not measure engagement. A cheaper check: train a model on OPEN and test it on a small set of older adults whose engagement is measured by an independent instrument; chance-level accuracy would falsify the dataset's core premise.

Watch

Extended reading notes

Core claim

The paper's central discovery is that a privacy-preserving dataset built only from extracted landmarks and behavioral features can support engagement recognition in older adult patients at accuracies comparable to those reported for student datasets. Using the HELP annotation protocol, three raters marked affective states (Bored, Calm/Satisfied, Confused/Frustrated, Motivated/Excited) and behavioral states (On-Task, Off-Task) second by second, and the binary engagement label was derived by the rule that Off-Task or Bored implies Not-Engaged. The dataset includes adaptive variable-length segments plus fixed 5-, 10-, and 30-second versions, and separates engagement detection (current state from current data) from engagement prediction (future state from preceding data), which the paper presents as a first in this literature. The best baseline, ROCKET on 3D facial landmarks, reaches 81.12% accuracy in 11-fold cross-validation for detection and 70.34% for prediction.

Load-bearing premise

The paper's labels treat visible affective and behavioral cues as a complete proxy for engagement, because annotators watched video without audio and therefore could not judge cognitive engagement, so if visible cues do not track older adults' real engagement, the models are learning to mimic the annotation rule rather than engagement itself.

Editorial extensions

If this is right

  • If OPEN supports reliable engagement detection, clinicians could monitor engagement in virtual rehab sessions in near real time without watching raw video.
  • The detection-versus-prediction distinction means future models can be trained not only to read current engagement but to forecast disengagement before it happens.
  • The longitudinal structure across six weekly sessions opens the door to studying how engagement evolves over a rehabilitation program, not just within a single class.
  • The landmark-only release provides a privacy-preserving template that other telehealth studies could adopt to share behavioral data.
  • The public availability of a benchmark for older adults invites direct comparison of algorithms across age groups, exposing where student-trained models fail.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own results suggest that the affective component is easier to detect (94.62% accuracy) than the behavioral component, which drives final engagement accuracy; a natural extension is to weight behavioral cues more heavily or to collect audio to improve behavioral labels.
  • Because annotations were based on video alone, the dataset's labels may encode visible expressiveness rather than true cognitive engagement; an external validation against self-report or clinician judgment would test whether the HELP rule transfers to this population.
  • The context-type annotations (group versus individual address) were collected but not used in the baselines; using them in a context-aware model is an obvious next step that the data already supports.
  • Cross-dataset transfer experiments, comparing OPEN-trained models on student datasets and vice versa, would quantify how much engagement cues differ with age and health status.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces OPEN, a dataset for automated engagement recognition in older adult patients during virtual cardiac rehabilitation learning sessions. The dataset includes over 35 hours of video-derived facial, hand, and body landmarks plus affect/behavior features, withheld raw video for privacy, second-level annotations from three raters (affective and behavioral components, binary engagement labels, and context types), and multiple fixed/variable-length sample versions. The authors also benchmark several machine-learning and deep-learning models for engagement detection and prediction, reporting up to 81% detection accuracy and 70% prediction accuracy.

Significance. If the dataset is used responsibly, it is a potentially valuable community resource: it addresses an under-studied population (older adult patients), provides a privacy-preserving landmark/feature format, reports inter-rater reliability values (Fleiss' kappa 0.81 for engagement), and is the largest public dataset of its kind. The authors are transparent about many limitations, including the small participant count, low demographic diversity, and absence of audio/raw video. However, the central claim that OPEN enables multi-component engagement recognition is undermined by the near-collinearity of the binary engagement label with the behavioral On-Task/Off-Task label, and by the fact that cognitive engagement was not annotated. The paper's headline accuracy figures also need to be interpreted alongside the much lower leave-one-participant-out results. These issues are substantive but addressable with reframing and additional analyses.

major comments (3)
  1. [§3.5.2, §3.6.5, Table 3, §4.3, Table 7] The binary engagement label is almost a recode of the behavioral On-Task/Off-Task label. Section 3.5.2 defines Not-Engaged iff Off-Task or (On-Task and Bored); Table 3 shows that Bored constitutes only 113 of 4,494 variable-length samples (2.5%). Thus the affective component changes the behavioral decision in a tiny fraction of cases. Table 7 confirms this empirically: behavioral detection accuracy (0.6052 LOPO, 0.8180 11-fold) is nearly identical to overall engagement accuracy (0.6059, 0.8210). Combined with the explicit statement in §3.5.2 that cognitive engagement was not annotated, the abstract's claim that OPEN supports engagement recognition based on affective, behavioral, and cognitive components is not supported by the released binary label. The paper should either rename the binary label (e.g., 'behaviorally engaged/on-task') or demonstrate that the affective component contributes meaningfully beyond the HELP rule.
  2. [§4 (Tables 4–6), abstract, conclusion] The headline '81% accuracy' is reported from 11-fold CV, in which the same participant can appear in both training and validation sets, as the paper itself notes at the end of Section 4. The more honest leave-one-participant-out accuracy for the same setting is about 0.64 (Table 5, 10 s / 8 fps ROCKET: 0.6362; Table 4, best LOPO Transformer: 0.6671). The paper should present LOPO as the primary generalization metric, report confidence intervals or per-fold variability, and temper the abstract and conclusion accordingly.
  3. [§4.2, Table 6] The engagement prediction task assigns to each sample the label of the immediately following sample from the same session. This is a next-sample (lag-1) classification within a session, not a longitudinal or long-horizon prediction. In 11-fold CV, the same participant's samples are in both training and testing, so the reported 0.7034 accuracy partly reflects participant-specific behavioral style; LOPO results are 0.5908 and 0.5686, barely above the engaged-class base rate (≈0.576 for 10 s samples). The claim of 'systematically investigating engagement prediction' should be qualified by these limitations.
minor comments (5)
  1. [§3.6.5 / Figure 2] The text refers to 'Figures 1(a) and (b)' when comparing behavioral and affective state change rates; the intended reference appears to be Figure 2(a) and (b).
  2. [§4.1] Table 4 does not report the frame rate, yet the text states that all results in Table 4 were obtained at fps=8. Please add a note or column header for clarity.
  3. [§4.1] The sentence 'The performance of the models and feature sets that achieved the highest results on OPEN with 10-second data samples in Table 5...' references Table 5, but the corresponding highest results appear in Table 4. Please correct the cross-reference.
  4. [Data Availability] The dataset is described as 'publicly available,' but access requires contacting the PI and executing a data-sharing agreement. Please use a term such as 'available upon request under a data-sharing agreement' to avoid ambiguity.
  5. [§3.6.6] Inter-rater reliability is reported for the three-state behavioral, five-state affective, and three-state engagement annotations. It would be useful to also report reliability for the final binary engagement label after collapsing Not-Visible, since this is the label used in the headline experiments.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: benchmark results are evaluated on held-out folds; the HELP-derived engagement label is transparent, and self-citations are contextual rather than load-bearing.

full rationale

This paper contributes a dataset and baseline benchmarks rather than a first-principles derivation, so there is no derivation chain whose output equals its input by construction. The engagement ground truth is produced by three annotators applying the external HELP protocol [28], with the binary label defined by a stated deterministic rule from affective and behavioral components (Section 3.5.2). Baseline detection and prediction models are assessed with 11-fold and leave-one-participant-out cross-validation on held-out samples, so the reported 81% detection accuracy is an empirical result, not a quantity forced by the label definition. The engagement prediction task (Section 4.2) targets the label of the following 10-second sample in the same session; the target is a later time step, not a re-expression of the input features, and no parameter is fitted to the future label and then renamed a prediction. The near-agreement between engagement and behavioral accuracies in Table 7 is disclosed by the authors and follows from the HELP rule together with the rarity of Bored annotations (Section 3.6.5); this is a construct-validity caveat, not a circular step. Self-citations such as [16] and [24] support framing, feature choices, and comparison points, but they are not load-bearing for the dataset's validity or for the benchmark numbers. The paper explicitly acknowledges in Section 3.5.2 that cognitive engagement could not be annotated because video was the sole modality; this limitation tempers any multi-component claim but does not make the evaluation circular. Overall, no identified step reduces to its own input, so the circularity score is low.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No new physical entities are postulated. The central claims rest on annotation validity, feature extractor reliability, and modeling choices. No derivation or external benchmark is used, so the ledger is dominated by domain assumptions about the ground truth and the feature pipeline.

free parameters (5)
  • Input frame rate (fps) = 8
    Video recorded at 16 fps; Table 5 shows 16 fps lowers LOPO accuracy and raises 11-fold accuracy, so this choice moves the headline number.
  • Sample duration for headline result = 10 seconds
    Ten-second clips are the literature standard and give the reported 81% detection; 5s gives 0.8347 and 30s 0.7556 in 11-fold CV.
  • Fixed-length labeling strategy = Strategy 1 (majority voting)
    Strategy 2 increases Off-Task labels and changes class distributions; the main results use Strategy 1.
  • Model and feature selection = ROCKET with 3D facial landmarks
    The best combination from many models and feature sets is highlighted in the abstract and conclusion.
  • ROCKET kernel count = 10000
    Chosen by the authors; no ablation or sensitivity analysis is reported.
assumptions (4)
  • domain assumption HELP rule-based combination (Off-Task or Bored implies Not-Engaged) gives valid binary engagement ground truth for older adult cardiac patients.
    Applied in Section 3.5.2 without population-specific validation; the rule originates from student-learning literature (Woolf et al. [83]).
  • domain assumption Cognitive engagement can be omitted when deriving the engagement label.
    Section 3.5.2 states annotators could annotate only affective and behavioral components because video was the sole data modality.
  • domain assumption OpenFace, EmoFAN, and MediaPipe produce reliable landmarks and features on 720p, 16 fps webcam video of older adults.
    Section 3.7 lists the extracted features; no failure analysis or quality control for landmark tracking is reported.
  • domain assumption High inter-rater agreement implies the labels are accurate ground truth.
    Section 3.6.6 reports Kappa values; agreement measures consistency, not construct validity against self-report or clinical outcomes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OPEN: A Benchmark Dataset and Baseline for Older Adult Patient Engagement Recognition in Virtual Rehabilitation Learning Environments." pith.science (2026). https://pith.science/paper/RFKRSIFP

@misc{pith2026250717959,
  author       = {Pith},
  title        = {Pith review of: OPEN: A Benchmark Dataset and Baseline for Older Adult Patient Engagement Recognition in Virtual Rehabilitation Learning Environments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RFKRSIFP}},
  note         = {Machine review of arXiv:2507.17959}
}
read the original abstract

Engagement in virtual learning is essential for participant satisfaction, performance, and adherence, particularly in online education and virtual rehabilitation, where interactive communication plays a key role. Yet, accurately measuring engagement in virtual group settings remains a challenge. There is increasing interest in using artificial intelligence (AI) for large-scale, real-world, automated engagement recognition. While engagement has been widely studied in younger academic populations, research and datasets focused on older adults in virtual and telehealth learning settings remain limited. Existing methods often neglect contextual relevance and the longitudinal nature of engagement across sessions. This paper introduces OPEN (Older adult Patient ENgagement), a novel dataset supporting AI-driven engagement recognition. It was collected from eleven older adults participating in weekly virtual group learning sessions over six weeks as part of cardiac rehabilitation, producing over 35 hours of data, making it the largest dataset of its kind. To protect privacy, raw video is withheld; instead, the released data include facial, hand, and body joint landmarks, along with affective and behavioral features extracted from video. Annotations include binary engagement states, affective and behavioral labels, and context-type indicators, such as whether the instructor addressed the group or an individual. The dataset offers versions with 5-, 10-, 30-second, and variable-length samples. To demonstrate utility, multiple machine learning and deep learning models were trained, achieving engagement recognition accuracy of up to 81 percent. OPEN provides a scalable foundation for personalized engagement modeling in aging populations and contributes to broader engagement recognition research.

Figures

Figures reproduced from arXiv: 2507.17959 by the authors.

Figure 1
Figure 1. Screenshot of a virtual cardiac rehabilitation educational session [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The frequency of data samples across various ranges of sample [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Distribution of the durations of the seven context types, as [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 77 canonical work pages

  1. [1]

    Rehabilita- tion,

    World Health Organization, “Rehabilita- tion,” https://www.who.int/news-room/fact- sheets/detail/rehabilitation, 2023, Accessed: January 30, 2023

  2. [2]

    Psychometric validation of the cardiac rehabilitation barriers scale,

    S. Shanmugasegaram, L. Gagliese, P . Oh, D. E. Stewart, S. J. Brister, V . Chan, and S. L. Grace, “Psychometric validation of the cardiac rehabilitation barriers scale,” Clinical rehabilitation, vol. 26, no. 2, pp. 152–164, 2012

  3. [3]

    Home-based cardiac rehabilitation among patients unwilling to participate in hospital- based programs,

    I. Nabutovsky, D. Breitner, A. Heller, Y. Levine, M. Moreno, M. Scheinowitz, C. Levin, and R. Klempfner, “Home-based cardiac rehabilitation among patients unwilling to participate in hospital- based programs,” Journal of Cardiopulmonary Rehabilitation and Prevention, vol. 44, no. 1, pp. 33–39, 2024

  4. [4]

    Wearable sensors and machine learning in post-stroke rehabilitation assessment: A systematic review,

    I. Boukhennoufa, X. Zhai, V . Utti, J. Jackson, and K. D. McDonald- Maier, “Wearable sensors and machine learning in post-stroke rehabilitation assessment: A systematic review,” Biomedical Signal Processing and Control, vol. 71, p. 103197, 2022

  5. [5]

    Telerehabilitation for patients with knee osteoarthritis: a focused review of technologies and teleservices,

    M. Naeemabadi, H. Fazlali, S. Najafi, B. Dinesen, J. Hansen et al., “Telerehabilitation for patients with knee osteoarthritis: a focused review of technologies and teleservices,”JMIR Biomedical Engineer- ing, vol. 5, no. 1, p. e16991, 2020

  6. [6]

    Ai-driven stroke rehabilitation systems and assessment: A systematic review,

    S. Rahman, S. Sarker, A. N. Haque, M. M. Uttsha, M. F. Islam, and S. Deb, “Ai-driven stroke rehabilitation systems and assessment: A systematic review,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2022

  7. [7]

    Artificial intelligence-driven virtual rehabilitation for people living in the community: A scoping review,

    A. Abedi, T. J. Colella, M. Pakosh, and S. S. Khan, “Artificial intelligence-driven virtual rehabilitation for people living in the community: A scoping review,” NPJ Digital Medicine, vol. 7, no. 1, p. 25, 2024. IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, VOL. 18, NO. 18, JUL Y 2025 13

  8. [8]

    A conceptual review of engagement in healthcare and rehabilita- tion,

    F. A. Bright, N. M. Kayes, L. Worrall, and K. M. McPherson, “A conceptual review of engagement in healthcare and rehabilita- tion,” Disability and rehabilitation, vol. 37, no. 8, pp. 643–654, 2015

Show all 89 references
  1. [9]

    Fa- cilitating neurorehabilitation through principles of engagement,

    M. M. Danzl, N. M. Etter, R. D. Andreatta, and P . H. Kitzman, “Fa- cilitating neurorehabilitation through principles of engagement,” Journal of allied health, vol. 41, no. 1, pp. 35–41, 2012

  2. [10]

    Unpacking patient engagement in remote consultation,

    Z. Liu, A. Brandon-Jones, and C. Vasilakis, “Unpacking patient engagement in remote consultation,” International Journal of Oper- ations & Production Management, vol. 44, no. 13, pp. 157–194, 2024

  3. [11]

    School engage- ment: Potential of the concept, state of the evidence,

    J. A. Fredricks, P . C. Blumenfeld, and A. H. Paris, “School engage- ment: Potential of the concept, state of the evidence,” Review of educational research, vol. 74, no. 1, pp. 59–109, 2004

  4. [12]

    The challenges of defining and measuring student engagement in science,

    G. M. Sinatra, B. C. Heddy, and D. Lombardi, “The challenges of defining and measuring student engagement in science,” pp. 1–13, 2015

  5. [13]

    Automatic context-aware inference of engagement in hmi: A survey,

    H. Salam, O. Celiktutan, H. Gunes, and M. Chetouani, “Automatic context-aware inference of engagement in hmi: A survey,” IEEE Transactions on Affective Computing, vol. 15, no. 2, pp. 445–464, 2024

  6. [14]

    Engagement detection and its applications in learning: a tutorial and selective review,

    B. M. Booth, N. Bosch, and S. K. D’Mello, “Engagement detection and its applications in learning: a tutorial and selective review,” Proceedings of the IEEE, vol. 111, no. 10, pp. 1398–1422, 2023

  7. [15]

    Engagement detection in online learning: a review,

    M. Dewan, M. Murshed, and F. Lin, “Engagement detection in online learning: a review,” Smart Learning Environments , vol. 6, no. 1, pp. 1–20, 2019

  8. [16]

    Inconsistencies in measuring student engagement in virtual learning–a critical review,

    S. S. Khan, A. Abedi, and T. Colella, “Inconsistencies in measuring student engagement in virtual learning–a critical review,” arXiv preprint arXiv:2208.04548, 2022

  9. [17]

    Automatic student engagement measurement using machine learning techniques: A literature study of data and methods,

    S. Mandia, R. Mitharwal, and K. Singh, “Automatic student engagement measurement using machine learning techniques: A literature study of data and methods,” Multimedia Tools and Applications, vol. 83, no. 16, pp. 49 641–49 672, 2024

  10. [18]

    Predicting en- gagement of older people’s virtual teams from video call analysis,

    N. Noceti, S. Campisi, A. Chirico, V . Cuculo, G. Grossi, M. Mich- elotto, F. Odone, A. Gaggioli, and R. Lanzarotti, “Predicting en- gagement of older people’s virtual teams from video call analysis,” International Journal of Human–Computer Interaction, pp. 1–12, 2024

  11. [19]

    Facial expression recognition influ- enced by human aging,

    G. Guo, R. Guo, and X. Li, “Facial expression recognition influ- enced by human aging,” IEEE Transactions on Affective Computing, vol. 4, no. 3, pp. 291–298, 2013

  12. [20]

    Age and gender differences in emotion recognition,

    L. Abbruzzese, N. Magnani, I. H. Robertson, and M. Mancuso, “Age and gender differences in emotion recognition,” Frontiers in psychology, vol. 10, p. 2371, 2019

  13. [21]

    Selective attention and facial expression recognition in patients with parkinson’s disease,

    L. Alonso-Recio, J. M. Serrano, and P . Mart ´ın, “Selective attention and facial expression recognition in patients with parkinson’s disease,” Archives of clinical neuropsychology, vol. 29, no. 4, pp. 374– 384, 2014

  14. [22]

    Y. Zhou, W. Han, X. Yao, J. Xue, Z. Li, and Y. Li, “Developing a machine learning model for detecting depression, anxiety, and apathy in older adults with mild cognitive impairment using speech and facial expressions: A cross-sectional observational study,” International journ...

  15. [23]

    Automatic engagement estima- tion in smart education/learning settings: a systematic review of engagement definitions, datasets, and methods,

    S. N. Karimah and S. Hasegawa, “Automatic engagement estima- tion in smart education/learning settings: a systematic review of engagement definitions, datasets, and methods,” Smart Learning Environments, vol. 9, no. 1, pp. 1–48, 2022

  16. [24]

    Engagement measurement based on facial landmarks and spatial-temporal graph convolutional net- works,

    A. Abedi and S. S. Khan, “Engagement measurement based on facial landmarks and spatial-temporal graph convolutional net- works,” in International Conference on Pattern Recognition. Springer, 2024, pp. 321–338

  17. [25]

    Tcct- net: Two-stream network architecture for fast and efficient engage- ment estimation via behavioral feature signals,

    A. Vedernikov, P . Kumar, H. Chen, T. Sepp ¨anen, and X. Li, “Tcct- net: Two-stream network architecture for fast and efficient engage- ment estimation via behavioral feature signals,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, ...

  18. [26]

    Do i have your attention: A large scale engagement prediction dataset and baselines,

    M. Singh, X. Hoque, D. Zeng, Y. Wang, K. Ikeda, and A. Dhall, “Do i have your attention: A large scale engagement prediction dataset and baselines,” in Proceedings of the 25th International Conference on Multimodal Interaction , ser. ICMI ’23. New York, NY, USA: Association fo...

  19. [27]

    On the influence of an iterative affect annotation approach on inter-observer and self-observer reliability,

    S. K. D’Mello, “On the influence of an iterative affect annotation approach on inter-observer and self-observer reliability,” IEEE Transactions on Affective Computing, vol. 7, no. 2, pp. 136–149, 2015

  20. [28]

    Human expert labeling process (help): towards a reliable higher-order user state labeling process and tool to assess student engagement,

    S. Aslan, S. E. Mete, E. Okur, E. Oktay, N. Alyuz, U. E. Genc, D. Stanhill, and A. A. Esme, “Human expert labeling process (help): towards a reliable higher-order user state labeling process and tool to assess student engagement,” Educational Technology, pp. 53–59, 2017

  21. [29]

    Baker rodrigo ocumpaugh monitoring protocol (bromp) 2.0 technical and training manual,

    J. Ocumpaugh, “Baker rodrigo ocumpaugh monitoring protocol (bromp) 2.0 technical and training manual,” New York, NY and Manila, Philippines: Teachers College, Columbia University and Ateneo Laboratory for the Learning Sciences, vol. 60, 2015

  22. [30]

    An unobtrusive and multimodal approach for behavioral engagement detection of students,

    N. Alyuz, E. Okur, U. Genc, S. Aslan, C. Tanriover, and A. A. Esme, “An unobtrusive and multimodal approach for behavioral engagement detection of students,” in Proceedings of the 1st ACM SIGCHI International Workshop on Multimodal Interaction for Educa- tion, 2017, pp. 26–32

  23. [31]

    Annotating student engagement across grades 1–12: Associations with demographics and expressivity,

    N. Alyuz, S. Aslan, S. K. D’Mello, L. Nachman, and A. A. Esme, “Annotating student engagement across grades 1–12: Associations with demographics and expressivity,” inInternational Conference on Artificial Intelligence in Education. Springer, 2021, pp. 42–51

  24. [32]

    Behavioral engagement detection of students in the wild,

    E. Okur, N. Alyuz, S. Aslan, U. Genc, C. Tanriover, and A. Ar- slan Esme, “Behavioral engagement detection of students in the wild,” in International Conference on Artificial Intelligence in Educa- tion. Springer, 2017, pp. 250–261

  25. [33]

    The faces of engagement: Automatic recognition of student en- gagementfrom facial expressions,

    J. Whitehill, Z. Serpell, Y.-C. Lin, A. Foster, and J. R. Movellan, “The faces of engagement: Automatic recognition of student en- gagementfrom facial expressions,” IEEE Transactions on Affective Computing, vol. 5, no. 1, pp. 86–98, 2014

  26. [34]

    Learner engagement measurement and classification in 1: 1 learning,

    S. Aslan, Z. Cataltepe, I. Diner, O. Dundar, A. A. Esme, R. Ferens, G. Kamhi, E. Oktay, C. Soysal, and M. Yener, “Learner engagement measurement and classification in 1: 1 learning,” in 2014 13th International Conference on Machine Learning and Applications. IEEE, 2014, pp. 545–552

  27. [35]

    Detecting student engagement: human versus ma- chine,

    N. Bosch, “Detecting student engagement: human versus ma- chine,” in Proceedings of the 2016 Conference on User Modeling Adaptation and Personalization, 2016, pp. 317–320

  28. [36]

    A hybrid intelligence-aided approach to affect-sensitive e-learning,

    J. Chen, N. Luo, Y. Liu, L. Liu, K. Zhang, and J. Kolodziej, “A hybrid intelligence-aided approach to affect-sensitive e-learning,” Computing, vol. 98, no. 1, pp. 215–233, 2016

  29. [37]

    Daisee: Towards user engagement recognition in the wild,

    A. Gupta, A. D’Cunha, K. Awasthi, and V . Balasubramanian, “Daisee: Towards user engagement recognition in the wild,” arXiv preprint arXiv:1609.01885, 2016

  30. [38]

    A crowdsourced approach to student engagement recognition in e-learning en- vironments,

    A. Kamath, A. Biswas, and V . Balasubramanian, “A crowdsourced approach to student engagement recognition in e-learning en- vironments,” in 2016 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 2016, pp. 1–9

  31. [39]

    Toward active and unobtrusive engagement assessment of distance learners,

    B. M. Booth, A. M. Ali, S. S. Narayanan, I. Bennett, and A. A. Farag, “Toward active and unobtrusive engagement assessment of distance learners,” in 2017 Seventh International Conference on Affective Computing and Intelligent Interaction (ACII) . IEEE, 2017, pp. 470–476

  32. [40]

    Predicting students’ attention in the classroom from kinect facial and body features,

    J. Zaletelj and A. Ko ˇsir, “Predicting students’ attention in the classroom from kinect facial and body features,” EURASIP journal on image and video processing, vol. 2017, no. 1, pp. 1–12, 2017

  33. [41]

    Prediction and localization of student engagement in the wild,

    A. Kaur, A. Mustafa, L. Mehta, and A. Dhall, “Prediction and localization of student engagement in the wild,” in 2018 Digital Image Computing: Techniques and Applications (DICTA). IEEE, 2018, pp. 1–8

  34. [42]

    Measuring student engagement level using facial infor- mation,

    I. Alkabbany, A. Ali, A. Farag, I. Bennett, M. Ghanoum, and A. Farag, “Measuring student engagement level using facial infor- mation,” in 2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019, pp. 3337–3341

  35. [43]

    Automatic recognition of student engagement using deep learning and facial expression,

    O. Mohamad Nezami, M. Dras, L. Hamey, D. Richards, S. Wan, and C. Paris, “Automatic recognition of student engagement using deep learning and facial expression,” in Joint European Confer- ence on Machine Learning and Knowledge Discovery in Databases . Springer, 2019, pp. 273–289

  36. [44]

    Application of deep learning on stu- dent engagement in e-learning environments,

    P . Bhardwaj, P . Gupta, H. Panwar, M. K. Siddiqui, R. Morales- Menendez, and A. Bhaik, “Application of deep learning on stu- dent engagement in e-learning environments,” Computers & Elec- trical Engineering, vol. 93, p. 107277, 2021

  37. [45]

    Student engagement dataset,

    K. Delgado, J. M. Origgi, T. Hasanpoor, H. Yu, D. Allessio, I. Ar- royo, W. Lee, M. Betke, B. Woolf, and S. A. Bargal, “Student engagement dataset,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 3628–3636

  38. [46]

    Hierarchical temporal multi- instance learning for video-based student learning engagement assessment

    J. Ma, X. Jiang, S. Xu, and X. Qin, “Hierarchical temporal multi- instance learning for video-based student learning engagement assessment.” in IJCAI, 2021, pp. 2782–2789

  39. [47]

    Facial emotion recog- nition based real-time learner engagement detection system in online learning context using deep learning models,

    S. Gupta, P . Kumar, and R. K. Tekchandani, “Facial emotion recog- nition based real-time learner engagement detection system in online learning context using deep learning models,” Multimedia Tools and Applications, pp. 1–30, 2022

  40. [48]

    Evaluation of e-learners’ concentra- IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, VOL. 18, NO. 18, JUL Y 2025 14 tion using recurrent neural networks,

    Y.-S. Jeong and N.-W. Cho, “Evaluation of e-learners’ concentra- IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, VOL. 18, NO. 18, JUL Y 2025 14 tion using recurrent neural networks,” The Journal of Supercomput- ing, pp. 1–18, 2022

  41. [49]

    Multi- label disengagement and behavior prediction in online learning,

    M. Verma, Y. Nakashima, N. Takemura, and H. Nagahara, “Multi- label disengagement and behavior prediction in online learning,” in International Conference on Artificial Intelligence in Education . Springer, 2022, pp. 633–639

  42. [50]

    Video-based affect detection in noninteractive learning environments

    Y. Chen, N. Bosch, and S. D’Mello, “Video-based affect detection in noninteractive learning environments.” International Educational Data Mining Society, 2015

  43. [51]

    Auto- mated detection of engagement using video-based estimation of facial expressions and heart rate,

    H. Monkaresi, N. Bosch, R. A. Calvo, and S. K. D’Mello, “Auto- mated detection of engagement using video-based estimation of facial expressions and heart rate,” IEEE Transactions on Affective Computing, vol. 8, no. 1, pp. 15–28, 2016

  44. [52]

    “en- gaged faces

    B. De Carolis, F. D’Errico, N. Macchiarulo, and G. Palestra, ““en- gaged faces”: Measuring and monitoring student engagement from face and gaze behavior,” in IEEE/WIC/ACM International Conference on Web Intelligence-Companion Volume, 2019, pp. 80–85

  45. [53]

    Time to scale: Gen- eralizable affect detection for tens of thousands of students across an entire school year,

    S. Hutt, J. F. Grafsgaard, and S. K. D’Mello, “Time to scale: Gen- eralizable affect detection for tens of thousands of students across an entire school year,” in Proceedings of the 2019 CHI conference on human factors in computing systems, 2019, pp. 1–14

  46. [54]

    Computer vision and human behaviour, emotion and cognition detection: A use case on student engagement,

    P . Vanneste, J. Oramas, T. Verelst, T. Tuytelaars, A. Raes, F. De- paepe, and W. Van den Noortgate, “Computer vision and human behaviour, emotion and cognition detection: A use case on student engagement,” Mathematics, vol. 9, no. 3, p. 287, 2021

  47. [55]

    Assessing student engagement from facial behavior in on-line learning,

    P . Buono, B. De Carolis, F. D’Errico, N. Macchiarulo, and G. Palestra, “Assessing student engagement from facial behavior in on-line learning,” Multimedia Tools and Applications , pp. 1–19, 2022

  48. [56]

    Au- tomatic prediction of presentation style and student engagement from videos,

    C. Thomas, K. P . Sarma, S. S. Gajula, and D. B. Jayagopi, “Au- tomatic prediction of presentation style and student engagement from videos,” Computers and Education: Artificial Intelligence , p. 100079, 2022

  49. [57]

    Using video to automatically detect learner affect in computer- enabled classrooms,

    N. Bosch, S. K. D’mello, J. Ocumpaugh, R. S. Baker, and V . Shute, “Using video to automatically detect learner affect in computer- enabled classrooms,” ACM Transactions on Interactive Intelligent Systems (TiiS), vol. 6, no. 2, pp. 1–26, 2016

  50. [58]

    Esti- mation of learners’ engagement using face and body features by transfer learning,

    X. Zheng, S. Hasegawa, M.-T. Tran, K. Ota, and T. Unoki, “Esti- mation of learners’ engagement using face and body features by transfer learning,” in International Conference on Human-Computer Interaction. Springer, 2021, pp. 541–552

  51. [59]

    A new emotion–based affective model to detect student’s engage- ment,

    K. Altuwairqi, S. K. Jarraya, A. Allinjawi, and M. Hammami, “A new emotion–based affective model to detect student’s engage- ment,” Journal of King Saud University-Computer and Information Sciences, vol. 33, no. 1, pp. 99–109, 2021

  52. [60]

    User engagement in an online digital health intervention to promote problem solving,

    H. L. O’Brien, A. T. Chen, J. Kaneshiro, and O. Zaslavsky, “User engagement in an online digital health intervention to promote problem solving,” Interacting with Computers , vol. 36, no. 5, pp. 355–369, 2024

  53. [61]

    Using response times to model student disengage- ment,

    J. E. Beck, “Using response times to model student disengage- ment,” in Proceedings of the ITS2004 Workshop on Social and Emo- tional Intelligence in Learning Environments , vol. 20, no. 2004. Ma- ceio, 2004, pp. 88–95

  54. [62]

    Improving state-of-the-art in detecting student engagement with resnet and tcn hybrid network,

    A. Abedi and S. S. Khan, “Improving state-of-the-art in detecting student engagement with resnet and tcn hybrid network,” in 2021 18th Conference on Robots and Vision (CRV) . IEEE, 2021, pp. 151– 157

  55. [63]

    Learning deep spatiotempo- ral feature for engagement recognition of online courses,

    L. Geng, M. Xu, Z. Wei, and X. Zhou, “Learning deep spatiotempo- ral feature for engagement recognition of online courses,” in 2019 IEEE Symposium Series on Computational Intelligence (SSCI) . IEEE, 2019, pp. 442–447

  56. [64]

    Class-attention video transformer for engagement intensity prediction,

    X. Ai, V . S. Sheng, and C. Li, “Class-attention video transformer for engagement intensity prediction,” arXiv preprint arXiv:2208.07216, 2022

  57. [65]

    Students engagement level detection in online e-learning using hybrid efficientnetb7 together with tcn, lstm, and bi-lstm,

    T. Selim, I. Elkabani, and M. A. Abdou, “Students engagement level detection in online e-learning using hybrid efficientnetb7 together with tcn, lstm, and bi-lstm,” IEEE Access , vol. 10, pp. 99 573–99 583, 2022

  58. [66]

    Predicting student engagement using sequential ensemble model,

    X. Tian, B. P . Nunes, Y. Liu, and R. Manrique, “Predicting student engagement using sequential ensemble model,” IEEE Transactions on Learning Technologies, 2023

  59. [67]

    Msc-trans: A multi-feature-fusion network with encoding structure for student engagement detecting,

    N. Xie, Z. Li, H. Lu, W. Pang, J. Song, and B. Lu, “Msc-trans: A multi-feature-fusion network with encoding structure for student engagement detecting,” IEEE Transactions on Learning Technologies, 2025

  60. [68]

    Openface 2.0: Facial behavior analysis toolkit,

    T. Baltrusaitis, A. Zadeh, Y. C. Lim, and L.-P . Morency, “Openface 2.0: Facial behavior analysis toolkit,” in2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018) . IEEE, 2018, pp. 59–66

  61. [69]

    Affect-driven ordinal engagement measurement from video,

    A. Abedi and S. S. Khan, “Affect-driven ordinal engagement measurement from video,” Multimedia Tools and Applications , pp. 1–20, 2023

  62. [70]

    Bag of states: A non-sequential approach to video-based engagement measurement,

    A. Abedi, C. Thomas, D. B. Jayagopi, and S. S. Khan, “Bag of states: A non-sequential approach to video-based engagement measurement,” arXiv preprint arXiv:2301.06730, 2023

  63. [71]

    Predicting engagement intensity in the wild using temporal convolutional network,

    C. Thomas, N. Nair, and D. B. Jayagopi, “Predicting engagement intensity in the wild using temporal convolutional network,” in Proceedings of the 20th ACM International Conference on Multimodal Interaction, 2018, pp. 604–610

  64. [72]

    Estimation of continuous valence and arousal levels from faces in naturalistic conditions,

    A. Toisoul, J. Kossaifi, A. Bulat, G. Tzimiropoulos, and M. Pantic, “Estimation of continuous valence and arousal levels from faces in naturalistic conditions,” Nature Machine Intelligence, vol. 3, no. 1, pp. 42–50, 2021

  65. [73]

    Marlin: Masked autoencoder for facial video representation learning,

    Z. Cai, S. Ghosh, K. Stefanov, A. Dhall, J. Cai, H. Rezatofighi, R. Haffari, and M. Hayat, “Marlin: Masked autoencoder for facial video representation learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 1493–1504

  66. [74]

    Detecting disengagement in virtual learning as an anomaly using temporal convolutional network autoencoder,

    A. Abedi and S. S. Khan, “Detecting disengagement in virtual learning as an anomaly using temporal convolutional network autoencoder,” Signal, Image and Video Processing, pp. 1–9, 2023

  67. [75]

    Me- diapipe: A framework for building perception pipelines,

    C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. Yong, J. Lee et al. , “Me- diapipe: A framework for building perception pipelines,” arXiv preprint arXiv:1906.08172, 2019

  68. [76]

    Spatial temporal graph convolutional networks for skeleton-based action recognition,

    S. Yan, Y. Xiong, and D. Lin, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018

  69. [77]

    Graph-based facial affect analysis: A review,

    Y. Liu, X. Zhang, Y. Li, J. Zhou, X. Li, and G. Zhao, “Graph-based facial affect analysis: A review,” IEEE Transactions on Affective Computing, 2022

  70. [78]

    canadienne de r ´eadaptation cardiaque, Canadian guidelines for cardiac rehabilitation and cardiovascular disease prevention: translating knowledge into action

    A. canadienne de r ´eadaptation cardiaque, Canadian guidelines for cardiac rehabilitation and cardiovascular disease prevention: translating knowledge into action . Canadian Association of Cardiac Rehabili- tation, 2009

  71. [79]

    Home-based versus centre-based cardiac rehabilitation,

    S. T. McDonagh, H. Dalal, S. Moore, C. E. Clark, S. G. Dean, K. Jolly, A. Cowie, J. Afzal, and R. S. Taylor, “Home-based versus centre-based cardiac rehabilitation,” Cochrane database of systematic reviews, no. 10, 2023

  72. [80]

    Clinical outcomes and qualitative per- ceptions of in-person, hybrid, and virtual cardiac rehabilitation,

    S. Ganeshan, H. Jackson, D. J. Grandis, D. Janke, M. L. Murray, V . Valle, and A. L. Beatty, “Clinical outcomes and qualitative per- ceptions of in-person, hybrid, and virtual cardiac rehabilitation,” Journal of cardiopulmonary rehabilitation and prevention, vol. 42, no. 5, pp...

  73. [81]

    Cardiac college - patient education program for cardiac rehabilitation,

    Health e-University, “Cardiac college - patient education program for cardiac rehabilitation,” n.d., accessed: 2025-01-04. [Online]. Available: https://www.healtheuniversity.ca/en/CardiacCollege

  74. [82]

    Revisiting annotations in online student engagement,

    S. Khan and S. Safa, “Revisiting annotations in online student engagement,” in Proceedings of the 2024 10th International Conference on Computing and Data Engineering, 2024, pp. 111–117

  75. [83]

    Affect-aware tutors: recognising and responding to student affect,

    B. Woolf, W. Burleson, I. Arroyo, T. Dragon, D. Cooper, and R. Picard, “Affect-aware tutors: recognising and responding to student affect,” International Journal of Learning Technology , vol. 4, no. 3/4, pp. 129–164, 2009

  76. [84]

    Dynamics of affective states during complex learning,

    S. D’Mello and A. Graesser, “Dynamics of affective states during complex learning,” Learning and Instruction, vol. 22, no. 2, pp. 145– 157, 2012

  77. [85]

    The dynamics of affective transitions in simulation problem-solving environments,

    R. S. d Baker, M. Rodrigo, T. Mercedes, and U. E. Xolocotzin, “The dynamics of affective transitions in simulation problem-solving environments,” in International Conference on Affective Computing and Intelligent Interaction. Springer, 2007, pp. 666–677

  78. [86]

    K. L. Gwet, Handbook of inter-rater reliability: The definitive guide to measuring the extent of agreement among raters . Advanced Analytics, LLC, 2014

  79. [87]

    Interrater reliability: the kappa statistic,

    M. L. McHugh, “Interrater reliability: the kappa statistic,” Bio- chemia medica, vol. 22, no. 3, pp. 276–282, 2012

  80. [88]

    Attention mesh: High-fidelity face mesh predic- tion in real-time,

    I. Grishchenko, A. Ablavatski, Y. Kartynnik, K. Raveendran, and M. Grundmann, “Attention mesh: High-fidelity face mesh predic- tion in real-time,” arXiv preprint arXiv:2006.10962, 2020

  81. [89]

    Rocket: exceptionally fast and accurate time series classification using random convolu- tional kernels,

    A. Dempster, F. Petitjean, and G. I. Webb, “Rocket: exceptionally fast and accurate time series classification using random convolu- tional kernels,” Data Mining and Knowledge Discovery, vol. 34, no. 5, pp. 1454–1495, 2020

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.