Pith. sign in

REVIEW 3 major objections 4 minor 76 references

TeachUp: Facilitating Early-Stage Teachers to Learn Instructional Strategies from Classroom Videos with Reflective Support

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read TeachUp, a system that structures video learning of teaching strategies into watch, practice, and evaluation stages, helps early-stage teachers apply what they learn from classroom videos to new lessons, outperforming a conventional…

desk verdict A worthwhile systems and benchmark paper whose main learning-effect claim rests on single-rater scores with no reliability evidence; good enough for peer review, but treat the p=.003 as provisional. read the letter →

arxiv 2608.08535 v1 pith:RXEX53NG submitted 2026-08-09 cs.HC

classification cs.HC
keywords video-basedlearninginstructionalstrategiesreflectivesupportmicroteachingLLMclassroomvideoanalysisteacherprofessionaldevelopmentwithin-subjectsstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that early-stage teachers learn instructional strategies from recorded classroom videos better when watching is wrapped in a structured reflective loop than when they watch and practice on their own. It presents TeachUp, which detects nine instructional-strategy categories in videos, prompts reflective questions while the video plays, generates a scenario-based microteaching task, and gives multimodal feedback comparing the user's performance with the original lesson. In a within-subjects study with 16 pre-service teachers, TeachUp users scored significantly higher on applying the learned strategy to a new teaching task (Z +0.32 vs -0.32, p = .003), and reported higher engagement, confidence, and intention to use the system than with the baseline. Interviews with four in-service teachers corroborate the usefulness of the strategy-clip detection and the reflective cycle, while also pointing to concerns about task flexibility and transfer to real classrooms.

What carries the argument

The central mechanism is the watch-practice-evaluate interaction loop, instantiated as a concrete implementation of an experiential learning cycle. The loop is powered by two computational components: a transcript-based detection pipeline that uses an ensemble of three large language models to segment classroom videos and tag clips with nine instructional-strategy categories, and a multimodal large language model that evaluates users' microteaching recordings and generates comparison-based feedback. The detection pipeline's rule-based aggregator merges overlapping predictions across models, and its ensemble strategy (accepting only clusters predicted by at least two models) is what yields the reported precision of 0.634.

What would settle it

Score the same set of 16 teaching-test videos with at least one additional independent rater blind to condition, using the same rubric, and compute inter-rater agreement; if agreement is low (intraclass correlation below roughly 0.6) or the TeachUp-versus-baseline difference on strategy application loses significance when rater identity is included in the model, the central claim would be undercut.

Watch

Extended reading notes

Core claim

The core claim is that the implicit pedagogical reasoning behind a teacher's classroom actions can be made learnable by pairing video examples with a reflection-and-practice loop. TeachUp operationalizes this as a watch-practice-evaluate cycle: an LLM-powered pipeline segments classroom videos into clips tagged with one of nine instructional strategies; reflective hints derived from a claim-evidence-reasoning-alternative protocol guide attention while watching; a scenario generator creates a microteaching task in which the user rehearses the strategy; and a multimodal LLM evaluates the practice video against the original instructional case. The experimental evidence compares TeachUp with a baseline that offers the same videos and free self-practice, and finds the reflective loop significantly improves expert-rated application of the strategy to a brand-new lesson, along with engagement and confidence. The paper presents the strategy-detection pipeline (precision 0.634 at IoU = 0.5) as an enabling component and releases a benchmark of 89 annotated strategy instances as a starting point for future work.

Load-bearing premise

The main learning-outcome result rests on a single expert rater's blind 7-point Likert scoring of each participant's short teaching-test video as a valid and reliable measure of applying learned teaching strategies, and the paper reports no inter-rater reliability or rubric-validation evidence for this measure.

Editorial extensions

If this is right

  • Early-stage teachers can transfer a strategy observed in a video to a new lesson more successfully when reflection and rehearsal are structured around the video, which suggests video libraries for teacher professional development should pair footage with interactive tasks rather than rely on self-directed viewing.
  • The transcript-based, LLM-driven detection pipeline offers a scalable way to index large collections of recorded classroom videos by instructional strategy, lowering the effort of finding relevant examples.
  • The significant gains in engagement and confidence imply that the reflective loop may reduce drop-off in self-paced online teacher learning, where motivation is a known barrier.
  • Because imitation of surface implementations showed no significant improvement, the effective mechanism appears to be understanding the strategy's timing and rationale rather than copying the observed teacher's actions.
  • The released benchmark of 89 annotated strategy instances provides a concrete test set for future work on automatic instructional-strategy detection in classroom videos.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the headline effect replicates under multi-rater scoring, the watch-practice-evaluate loop could be embedded into existing teacher-education platforms as a lightweight asynchronous layer, since the pipeline is designed as plug-and-play middleware.
  • The single-rater outcome measure leaves open the possibility that the reported effect size is inflated by scoring noise; a cheap replication with two independent raters and a pre-registered rubric is the natural next step.
  • Because the detection pipeline works from transcripts alone, strategies that are primarily non-verbal (e.g., nonlinguistic representations, cooperative-learning arrangements) may be under-detected, which would bias which strategies learners get offered; this is a testable implication the paper does not address.
  • The transfer-distance tension that experienced teachers raised suggests an adaptive version of the practice generator: adjust how far the simulated scenario departs from the observed video based on the user's demonstrated understanding, something the current system does not attempt.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents TeachUp, an interactive system that helps early-stage teachers learn instructional strategies from classroom videos through a three-stage 'watch-practice-evaluate' loop. The authors first report a formative study (N=9) that motivates design requirements, then describe an LLM-powered pipeline that detects nine Marzano instructional strategies in classroom videos (claimed precision 0.634 at IoU=0.5), followed by a within-subjects study (N=16) comparing TeachUp with a baseline video-watching and self-practicing system. The main quantitative claims are that TeachUp improves participants' application of learned strategies to a new teaching task (Z-scored expert rating, p=.003), increases engagement (Q4, p=.026), confidence (Q1, p=.009), perceived practice benefit (Q3, p=.014), and intention to use (Q6, p=.017), while other performance dimensions are marginal or nonsignificant. Interviews with four in-service teachers are used to generalize the findings.

Significance. If the central effectiveness claims hold, TeachUp would be a useful contribution to HCI and teacher professional development, since it directly addresses a known gap in video-based learning: helping novice teachers notice, rehearse, and reflect on context-dependent instructional strategies. The paper also contributes a released benchmark of 89 annotated strategy segments and a reproducible LLM-based detection pipeline, which are valuable starting points even with modest precision. The study design is thoughtful in several respects: the user study uses a within-subjects design with counterbalanced tasks, blind expert raters, a pre-experiment preparation phase, and an incentive structure, and the qualitative data are analyzed with a transparent thematic approach. The paper is also honest about several limitations (e.g., the absence of a component-level ablation and the lack of long-term behavioral outcomes). However, the primary learning-outcome measure rests on single-rater holistic judgments with no reliability evidence, and the pipeline evaluation is an in-sample fit, so the strongest empirical claims are not yet fully supported.

major comments (3)
  1. [§5.2.4, Table 1 (Dim2)] The central RQ1 result—that TeachUp improves application of learned instructional strategies (p=.003)—rests entirely on holistic 7-point Likert scores given by a single expert rater per participant, with no inter-rater reliability, rubric validation, or duplicate scoring. Section 5.2.4 states that one mathematics teacher rated math tests and one Chinese teacher rated Chinese tests; each participant therefore contributes one score from each of two different raters on two different subject tasks (Cues/Advance Organizers in math vs. Cooperative Learning in Chinese). The Z-score standardization in Section 6 normalizes only each rater's marginal distribution and cannot remove rater-by-condition or rater-by-task interactions. A holistic judgment of 'application' with no demonstrated reliability could be influenced by rater expectations, perceived fluency, or task-difficulty differences, any of which could produce or mask the reported effect. This is load-bearing because Dim2 is the primary performance claim, while Dim1, Dim3, and Dim4 are marginal or nonsignificant. I ask the authors to provide inter-rater reliability on at least a subset of double-scored test videos, a rubric that maps the four metrics to observable behaviors, and an analysis or explicit acknowledgment of the rater/task confound in the paired comparison (Section 8.4 does not currently address this).
  2. [§4.2.1 (Pipeline technical evaluation)] The claimed pipeline precision of 0.634 is measured on the same benchmark from which the few-shot examples were drawn. The text states that 'we incorporated transcript segments as positive few-shot examples and frequently misclassified segments as negative few-shot examples' and that these examples were designed using the benchmark videos; the precision is then reported on that same benchmark. This is an in-sample fit, not an out-of-sample estimate, so the reported precision is likely optimistic relative to what users would encounter on new classroom videos. The zero-shot comparison (precision=0.095) is useful but does not resolve the circularity. Please report a held-out or cross-validated evaluation, or clearly frame the reported number as a development-set fit. This matters because the pipeline is presented as an enabling contribution and because Section 7's interview claims about clip quality are based on pipeline outputs that were not independently error-checked.
  3. [§6, Table 1] The paper reports p-values for four expert-rated dimensions and six questionnaire items without any correction for multiple comparisons. Several effects are marginal (Dim1 p=.065, Dim4 p=.086, Q5 p=.053), and the headline effects (Dim2 p=.003, Q1 p=.009, Q3 p=.014, Q4 p=.026, Q6 p=.017) are drawn from the same set of participants across two conditions. Since the primary claims are selective, I ask the authors to report the number of comparisons considered, to apply or justify not applying a correction (e.g., Bonferroni or FDR), and to interpret marginal results accordingly.
minor comments (4)
  1. [§4.2.1] The arithmetic of the annotation agreement is unclear: the text reports 85 agreed events, 1 event added by one annotator, 4 events added by the other, and 5 disagreements resolved through discussion, which sums to 95 rather than the stated final benchmark of 89 instances. Please reconcile these counts.
  2. [§4.2.1] The sentence 'traditional inter-rater statistics such as Cohen's κ or ICC are not applicable' is too strong; segment-level agreement measures can be computed for temporal detection tasks, and the authors' IoU-based agreement already provides a reasonable alternative. A brief justification or reference would help.
  3. [§5.2.2] The two learning tasks differ in both subject and strategy, and the two tasks are assigned to conditions with counterbalancing only of task order and system order, not of task-condition pairing. Because each participant always does one task in one condition, any subject-level or strategy-level difficulty difference is fully confounded with condition in the paired analysis. This should at least be acknowledged in Section 8.4.
  4. [Appendix C] Questionnaire items Q1–Q6 are reported with p-values but no per-item effect sizes or confidence intervals; reporting these would help readers assess practical significance alongside the Wilcoxon tests.

Circularity Check

1 steps flagged · score 4.0 of 10

Technical pipeline precision is an in-sample estimate, but the central HCI user-study claim is independent and not circular.

  1. fitted input called prediction [Section 4.2.1, 'Detecting Instructional Strategies in Classroom Videos', 'Pipeline' and 'Technical evaluation' paragraphs]
    "To improve detection precision, we first applied a zero-shot prompt on eight carefully selected high-quality videos and analyzed the models’ typical errors... We then incorporated transcript segments as positive few-shot examples and frequently misclassified segments as negative few-shot examples. ... At IoU = 0.5—a common threshold in object detection [15]—the ensemble achieves a precision of 0.634 of detecting nine strategies."

    The same benchmark of nine videos supplies both the tuning data and the evaluation data. After selecting positive few-shot examples from 'eight carefully selected high-quality videos' and negative examples from the same corpus to shape the detector, the paper reports precision on that same benchmark. The 0.634 precision at IoU = 0.5 is therefore an in-sample fit to the data used to build the prompt, not an out-of-sample estimate of detection performance. The paper's own limitation section says the pipeline was 'developed and validated on a small benchmark of 89 video scripts,' and the controlled user study used manually annotated ground-truth clips (Section 8.4).

full rationale

The main derivation chain is the within-subjects user study: TeachUp versus a video-watching baseline, with blind expert ratings of a new teaching task as the outcome. That outcome is not derived from the pipeline's detections or from the system's own generated outputs; the paper explicitly states in Section 8.4 that 'the controlled study used manually annotated ground-truth clips,' so the central effectiveness claim is independent of the pipeline's in-sample precision. No self-citation is load-bearing: the citation to Wu et al. [60] merely inspires the verbal-reflection coding scheme and does not entail the observed differences. The one concrete circularity is the pipeline evaluation, where few-shot examples are selected from the same benchmark on which precision is measured, making the reported 0.634 precision partly a fit to the evaluation data. The single-rater, no-reliability expert scoring and the cross-rater/cross-task paired comparison are serious threats to the validity of the p=.003 learning-outcome result, but those are measurement-validity concerns rather than circularity, so they are not counted in the score.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claims rest on the pedagogical framework (Marzano, LAP, Kolb), on the accuracy of ground-truth annotations and LLM outputs, and on study measurements that are partly unvalidated. No new physical entities are introduced.

free parameters (4)
  • IoU matching threshold = 0.5
    Used in Section 4.2.1 to define whether a predicted segment matches ground truth; the headline precision of 0.634 is reported at this threshold.
  • Ensemble cluster overlap ratio = 0.5
    Section 4.2.1: predictions are merged if temporal overlap exceeds 0.5.
  • Minimum model count per accepted cluster = at least 2 of 3 LLMs
    Section 4.2.1: a cluster is accepted only if it contains predictions from at least two models; this hand-set rule affects the reported precision and recall.
  • Few-shot example counts = 2 to 4 per event type
    Section 4.2.1: 'For each event, we designed two to four few-shot examples', chosen after error analysis on eight videos from the benchmark.
assumptions (5)
  • domain assumption Marzano et al.'s nine instructional strategies are a valid and sufficient taxonomy for the target teaching behaviors.
    Section 3.3: the system is built around these nine categories. If this taxonomy is incomplete or not appropriate for the subjects and videos, the detection and reflection loop may miss key strategies.
  • domain assumption Reflective support in the form of questions, hints, and LLM-generated feedback improves learning of instructional strategies.
    The design is grounded in Kolb's cycle and the Lesson Analysis Protocol (Sections 2.3, 4.1, 8.2). The study tests this assumption but does not independently prove the underlying learning theory.
  • domain assumption The ground-truth annotations in the benchmark are correct and reliable.
    Section 4.2.1: two authors annotated the videos with 94.4% agreement. The pipeline's precision is measured against these labels, so annotation errors would affect the reported performance.
  • domain assumption The LLM outputs, including transcripts, strategy detections, reflective hints, and evaluations, are sufficiently accurate and unbiased for the learning task.
    The system relies on DeepSeek-v3, Kimi-k2, GPT-4.1, and Qwen-2.5-Omni. No automated validation of generated content is reported, and the controlled study used manually selected clips rather than pipeline output (Section 8.4).
  • domain assumption A single short microteaching demonstration after a learning session measures transfer of instructional strategies.
    Section 5.2.4: expert ratings of a 3-5 minute test teaching are used as the main learning outcome. No reliability or construct-validation evidence is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TeachUp: Facilitating Early-Stage Teachers to Learn Instructional Strategies from Classroom Videos with Reflective Support." pith.science (2026). https://pith.science/paper/RXEX53NG

@misc{pith2026260808535,
  author       = {Pith},
  title        = {Pith review of: TeachUp: Facilitating Early-Stage Teachers to Learn Instructional Strategies from Classroom Videos with Reflective Support},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RXEX53NG}},
  note         = {Machine review of arXiv:2608.08535}
}
read the original abstract

Recorded videos of offline open classes provide good examples for early-stage teachers to learn instructional strategies, e.g., how to organize cooperative learning. However, learning by watching these videos is challenging, as these strategies are implicitly performed, and it lacks in-situ reflective support. In this paper, via a formative study (N=9), we design TeachUp to support the learning of instructional strategies from classroom teaching videos. TeachUp adopts an LLM-powered pipeline to detect nine instructional strategies in videos (precision = 63.4%), provides reflective questions and hints while watching, and generates customized practices with reflective feedback. A within-subjects study (N=16) shows that compared to a traditional video-playing and self-practicing baseline, early-stage teachers with TeachUp are more engaged in learning and perform better in applying learned strategies to new tasks. Interviews with four in-service teachers further generalize our findings and TeachUp's use cases. We discuss practical implications for fostering video-based learning of instructional strategies.

Figures

Figures reproduced from arXiv: 2608.08535 by the authors.

Figure 1
Figure 1. TeachUp’s interface in the watching stage. The interface is originally in Chinese and translated by GPT-4.1 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. TeachUp’s interface in the practicing stage (A, B, C, D) and evaluating stage (A, B’, C’, D’, E). The interface is originally in Chinese and translated by GPT-4.1 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. The computational pipeline of detecting instruc [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The architecture of TeachUp. The Video Clip is the output of the pipeline introduced in Section 4.2.1. 4.2.2 Supporting User Interactions. As shown in [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The video-watching page of the prototype. Translated from Chinese. [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: The evaluation page of the prototype. Translated from Chinese. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

76 extracted references · 64 canonical work pages

  1. [1]

    Allen and K

    D.W. Allen and K. Ryan. 1969.Microteaching. Addison-Wesley Publishing Com- pany. https://books.google.com.sg/books?id=iL1RAQAAIAAJ

  2. [2]

    Abdurrahman Ghaleb Almekhlafi, Sadiq Abdulwahed Ismail, and Abdel- moniem Ahmed Hassan. 2020. Teachers’ Reported Use of Marzano’s Instructional Strategies in United Arab Emirates K-12 Schools.International Journal of Instruc- tion13, 1 (2020), 325–340

  3. [3]

    Julie M Amador, Jode Keehr, Abraham Wallin, and Christopher Chilton. 2020. Video complexity: Describing videos used for teacher learning.Eurasia Journal of Mathematics, Science and Technology Education16, 4 (2020), em1834

  4. [4]

    Riku Arakawa and Hiromu Yakura. 2020. INWARD: A computer-supported tool for video-reflection improves efficiency and effectiveness in executive coaching. InProceedings of the 2020 CHI conference on human factors in computing systems. 1–13

  5. [5]

    Riku Arakawa, Hiromu Yakura, and Masataka Goto. 2023. CatAlyst: domain- extensible intervention for preventing task procrastination using large generative models. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–19

  6. [6]

    Laura Baecher, Shiao-Chuan Kung, Sarah Laleman Ward, and Kimberly Kern

  7. [7]

    Meg Schleppenbach Bates, Lena Phalen, and Cheryl G Moran. 2016. If you build it, will they reflect? Examining teachers’ use of an online video-based learning website.Teaching and teacher education58 (2016), 17–27

  8. [8]

    Marit Bentvelzen, Paweł W Woźniak, Pia SF Herbes, Evropi Stefanidi, and Jasmin Niess. 2022. Revisiting reflection in hci: Four design resources for technologies that support reflection.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies6, 1 (2022), 1–27

Show all 76 references
  1. [9]

    Amanda Berry, Fien Depaepe, and Jan Van Driel. 2016. Pedagogical content knowledge in teacher education. InInternational handbook of teacher education: volume 1. Springer, 347–386

  2. [10]

    Cynthia E Bolt-Lee. 2021. Developments in Research-Based Instructional Strate- gies: Learning-Centered Approaches for Accounting Education.E-Journal of Business Education and Scholarship of Teaching15, 2 (2021), 1–14

  3. [11]

    2013.Reflection: Turning experience into learning

    David Boud, Rosemary Keogh, and David Walker. 2013.Reflection: Turning experience into learning. Routledge

  4. [12]

    Evelyn M Boyd and Ann W Fales. 1983. Reflective learning: Key to learning from experience.Journal of humanistic psychology23, 2 (1983), 99–117

  5. [13]

    Jiwon Chun, Yuling Zhuang, Armanto Sutedjo, Colin Xu, Rong Ren, and Meng Xia. 2026. ArguMath: AI-Simulated Environment for Pre-service Teacher Train- ing in Orchestrating Classroom Mathematics Argumentation. InInternational Conference on Artificial Intelligence in Education. S...

  6. [14]

    Vanessa Echeverria, Lixiang Yan, Linxuan Zhao, Sophie Abel, Riordan Alfredo, Samantha Dix, Hollie Jaggard, Rosie Wotherspoon, Abra Osborne, Simon Bucking- ham Shum, et al. 2024. TeamSlides: A multimodal teamwork analytics dashboard for teacher-guided reflection in a physical l...

  7. [15]

    Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010. The pascal visual object classes (voc) challenge.Inter- national journal of computer vision88, 2 (2010), 303–338

  8. [16]

    Rowanne Fleck and Geraldine Fitzpatrick. 2010. Reflecting on reflection: framing a design landscape. InProceedings of the 22nd conference of the computer-human interaction special interest group of australia on computer-human interaction. 216– 223

  9. [17]

    Cyrille Gaudin and Sébastien Chaliès. 2015. Video viewing in teacher education and professional development: A literature review.Educational research review 16 (2015), 41–67

  10. [18]

    2011.Applied thematic analysis

    Greg Guest, Kathleen M MacQueen, and Emily E Namey. 2011.Applied thematic analysis. sage publications

  11. [19]

    Björn Haßler, Sara Hennessy, and Riikka Hofmann. 2020. OER4Schools: Outcomes of a sustained professional development intervention in sub-Saharan Africa. In Frontiers in Education, Vol. 5. Frontiers Media SA, 146

  12. [20]

    Chuanjun He and Chunmei Yan. 2011. Exploring authenticity of microteaching in pre-service teacher education programmes.Teaching education22, 3 (2011), 291–302

  13. [21]

    Brad Hokanson and Simon Hooper. 2004. Levels of teaching: A taxonomy for instructional design.Educational technology44, 6 (2004), 14–22

  14. [22]

    Mohammed Hoque, Matthieu Courgeon, Jean-Claude Martin, Bilge Mutlu, and Rosalind W Picard. 2013. Mach: My automated conversation coach. InProceed- ings of the 2013 ACM international joint conference on Pervasive and ubiquitous computing. 697–706

  15. [23]

    Juho Kim, Phu Tran Nguyen, Sarah Weir, Philip J Guo, Robert C Miller, and Krzysztof Z Gajos. 2014. Crowdsourcing step-by-step information extraction to enhance existing how-to videos. InProceedings of the SIGCHI conference on human factors in computing systems. 4017–4026

  16. [24]

    Taewan Kim, Seolyeong Bae, Hyun Ah Kim, Su-woo Lee, Hwajung Hong, Chanmo Yang, and Young-Ho Kim. 2024. MindfulDiary: Harnessing large language model to support psychiatric patients’ journaling. InProceedings of the 2024 CHI Confer- ence on Human Factors in Computing Systems. 1–20

  17. [25]

    Youngmin Kim, Jiwan Chung, Jisoo Kim, Sunghyun Lee, Sangkyu Lee, Junhyeok Kim, Cheoljong Yang, and Youngjae Yu. 2025. Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video- Grounded Dialogues. InProceedings of the 63rd Annual Meeting...

  18. [26]

    Seth King, Joseph Boyer, Tyler Bell, and Anne Estapa. 2022. An automated virtual reality training system for teacher-student interaction: A randomized controlled trial.JMIR serious games10, 4 (2022), e41097

  19. [27]

    2006.Evaluating training programs: The four levels

    Donald Kirkpatrick and James Kirkpatrick. 2006.Evaluating training programs: The four levels. Berrett-Koehler Publishers

  20. [28]

    Rafal Kocielnik, Lillian Xiao, Daniel Avrahami, and Gary Hsieh. 2018. Reflection companion: a conversational system for engaging users in reflection on physical activity.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies2, 2 (2018), 1–26

  21. [29]

    2014.Experiential learning: Experience as the source of learning and development

    David A Kolb. 2014.Experiential learning: Experience as the source of learning and development. FT press

  22. [30]

    Serge Leblanc. 2018. Analysis of video-based training approaches and professional development.Contemporary Issues in Technology and Teacher Education18, 1 (2018), 125–148

  23. [31]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437(2024)

  24. [32]

    Jingyuan Liu, Nazmus Saquib, Chen Zhutian, Rubaiat Habib Kazi, Li-Yi Wei, Hongbo Fu, and Chiew-Lan Tai. 2022. Posecoach: A customizable analysis and visualization system for video-based running coaching.IEEE Transactions on Visualization and Computer Graphics30, 7 (2022), 3180–3195

  25. [33]

    Arnold M Lund. 2001. Measuring usability with the use questionnaire12.Usability interface8, 2 (2001), 3–6

  26. [34]

    Shuai Ma, Taichang Zhou, Fei Nie, and Xiaojuan Ma. 2022. Glancee: An adaptable system for instructors to grasp student learning status in synchronous online classes. InProceedings of the 2022 CHI conference on human factors in computing systems. 1–25

  27. [35]

    Ahmed Magooda, Diane Litman, Ahmed Ashraf, and Muhsin Menekse. 2022. Improving the quality of students’ written reflections using natural language processing: Model design and classroom evaluation. InInternational conference on artificial intelligence in education. Springer, 519–525

  28. [36]

    2003.What works in schools: Translating research into action

    Robert J Marzano. 2003.What works in schools: Translating research into action. Ascd

  29. [37]

    2001.A Handbook for Classroom Instruction That Works.ERIC

    Robert J Marzano, Jennifer S Norford, Diane E Paynter, Debra J Pickering, and Barbara B Gaddy. 2001.A Handbook for Classroom Instruction That Works.ERIC

  30. [38]

    2001.Classroom instruction that works: Research-based strategies for increasing student achievement

    Robert J Marzano, Debra Pickering, and Jane E Pollock. 2001.Classroom instruction that works: Research-based strategies for increasing student achievement. Ascd

  31. [39]

    Toni-Jan Keith Palma Monserrat, Yawen Li, Shengdong Zhao, and Xiang Cao

  32. [40]

    Tahmina Nazari, Floyd W van de Graaf, Mary EW Dankbaar, Johan F Lange, Jeroen JG van Merriënboer, and Theo Wiggers. 2020. One step at a time: step by step versus continuous video-based learning to prepare medical students for performing surgical procedures.Journal of surgical ...

  33. [41]

    Seyed Parsa Neshaei, Thiemo Wambsganss, Hind El Bouchrifi, and Tanja Käser

  34. [42]

    Tricia J Ngoon, S Sushil, Angela EB Stewart, Ung-Sang Lee, Saranya Venkatraman, Neil Thawani, Prasenjit Mitra, Sherice Clarke, John Zimmerman, and Amy Ogan

  35. [43]

    Sitong Pan, Robin Schmucker, Bernardo Garcia Bulle Bueno, Salome Aguilar Llanes, Fernanda Albo Alarcón, Hangxiao Zhu, Adam Teo, and Meng Xia. 2025. Tutorup: What if your students were simulated? training tutors to address en- gagement challenges in online learning. InProceedin...

  36. [44]

    2012.Using technology with classroom instruction that works

    Howard Pitler, Elizabeth R Hubbell, and Matt Kuhn. 2012.Using technology with classroom instruction that works. Ascd

  37. [45]

    Ambili Remesh. 2013. Microteaching, an efficient technique for learning effective teaching.Journal of research in medical sciences: the official journal of Isfahan University of Medical Sciences18, 2 (2013), 158

  38. [46]

    Peter J Rich and Michael Hannafin. 2009. Video annotation tools: Technologies to scaffold, structure, and transform teacher reflection.Journal of teacher education 60, 1 (2009), 52–67

  39. [47]

    Marija Sablić, Ana Mirosavljević, and Alma Škugor. 2021. Video-based learning (VBL)—past, present and future: An overview of the research published from 2008 to 2019.Technology, Knowledge and Learning26, 4 (2021), 1061–1077

  40. [48]

    Alpay Sabuncuoglu and T Metin Sezgin. 2023. Developing a multimodal classroom engagement analysis dashboard for higher-education.Proceedings of the ACM on Human-Computer Interaction7, EICS (2023), 1–23

  41. [49]

    Lee S Shulman. 1986. Those who understand: Knowledge growth in teaching. Educational researcher15, 2 (1986), 4–14

  42. [50]

    Christina Siry and Sonya N Martin. 2014. Facilitating reflexivity in preservice science teacher education using video analysis and cogenerative dialogue in field- based methods courses.EURASIA journal of mathematics, science and technology education10, 5 (2014), 481–508

  43. [51]

    Winnie Wing-mui So. 2012. Quality of learning outcomes in an online video-based learning community: Potential and challenges for student teachers.Asia-Pacific Journal of Teacher Education40, 2 (2012), 143–158

  44. [52]

    Rand J Spiro, Brian P Collins, and Aparna Ramchandran. 2014. Reflections on a Post-Gutenerg Epistemology for Video Use in Ill-Structured Domains: Fostering Complex Learning and Cognitive Flexibility. InVideo research in the learning sciences. Routledge, 93–100

  45. [53]

    2009.The teaching gap: Best ideas from the world’s teachers for improving education in the classroom

    James W Stigler and James Hiebert. 2009.The teaching gap: Best ideas from the world’s teachers for improving education in the classroom. Simon and Schuster

  46. [54]

    Joseph A Taylor, Kathleen Roth, Christopher D Wilson, Molly AM Stuhlsatz, and Elizabeth Tipton. 2017. The effect of an analysis-of-practice, videocase-based, teacher professional development program on elementary students’ science achievement.Journal of Research on Educational...

  47. [55]

    Pugazhenthan Thangaraju and Bikash Medhi. 2023. Microteaching: Overview and examination evaluation.Indian Journal of Pharmacology55, 4 (2023), 257–262

  48. [56]

    Meredith Thompson, Kesiena Owho-Ovuakporie, Kevin Robinson, Yoon Jeon Kim, Rachel Slama, and Justin Reich. 2019. Teacher Moments: A digital simulation for preservice teachers to approximate parent–teacher conversations.Journal of Digital Learning in Teacher Education35, 3 (201...

  49. [57]

    Isabell Tucholka and Bernadette Gold. 2025. Analysing classroom videos in teacher education—How different instructional settings promote student teachers’ professional vision of classroom management.Learning and Instruction97 (2025), 102084

  50. [58]

    Xu Wang, Meredith Thompson, Kexin Yang, Dan Roy, Kenneth R Koedinger, Car- olyn P Rose, and Justin Reich. 2021. Practice-based teacher questioning strategy training with ELK: A role-playing simulation for eliciting learner knowledge. Proceedings of the ACM on Human-Computer In...

  51. [59]

    Xingbo Wang, Haipeng Zeng, Yong Wang, Aoyu Wu, Zhida Sun, Xiaojuan Ma, and Huamin Qu. 2020. Voicecoach: Interactive evidence-based training for voice modulation skills in public speaking. InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1–12

  52. [60]

    Shiwei Wu, Mingxiang Wang, Chuhan Shi, and Zhenhui Peng. 2025. ComViewer: An Interactive Visual Tool to Help Viewers Seek Social Support in Online Mental Health Communities.Proceedings of the ACM on Human-Computer Interaction9, 2 (2025), 1–31. TeachUp UIST ’26, November 02–05,...

  53. [61]

    Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, et al. 2025. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215(2025)

  54. [62]

    Tao Xu, Yuan Liu, Yaru Jin, Yueyao Qu, Jie Bai, Wenlan Zhang, and Yun Zhou

  55. [63]

    Xiaotong Xu, Jiayu Yin, Catherine Gu, Jenny Mar, Sydney Zhang, Jane L E, and Steven P Dow. 2024. Jamplate: Exploring llm-enhanced templates for idea reflection. InProceedings of the 29th International Conference on Intelligent User Interfaces. 907–921

  56. [64]

    Haipeng Zeng, Xinhuan Shu, Yanbang Wang, Yong Wang, Liguo Zhang, Ting- Chuen Pong, and Huamin Qu. 2020. Emotioncues: Emotion-oriented visual summarization of classroom videos.IEEE transactions on visualization and com- puter graphics27, 7 (2020), 3168–3181

  57. [65]

    Chao Zhang, Kexin Ju, Peter Bidoshi, Yu-Chun Grace Yen, and Jeffrey M Rzes- zotarski. 2025. Friction: Deciphering Writing Feedback into Writing Revisions through LLM-Assisted Reflection. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–27

  58. [66]

    Gefei Zhang, Shenming Ji, Yicao Li, Jingwei Tang, Jihong Ding, Meng Xia, Guodao Sun, and Ronghua Liang. 2025. CPVis: Evidence-based Multimodal Learning Analytics for Evaluation in Collaborative Programming. InProceedings of the 2025 CHI Conference on Human Factors in Computing...

  59. [67]

    From recorded to AI-generated instructional videos: A comparison of learning performance and experience.British Journal of Educational Technology 56, 4 (2025), 1463–1487

  60. [68]

    Meilan Zhang, Mary Lundeberg, Tom J McConnell, Matthew J Koehler, and Jan Eberhardt. 2010. Using questioning to facilitate discussion of science teach- ing problems in teacher professional development.Interdisciplinary Journal of Problem-Based Learning4, 1 (2010), 5

  61. [69]

    Zheyuan Zhang, Daniel Zhang-Li, Jifan Yu, Linlu Gong, Jinchang Zhou, Zhanxin Hao, Jianxiao Jiang, Jie Cao, Huiqin Liu, Zhiyuan Liu, et al . 2025. Simulating classroom education with llm-empowered agents. InProceedings of the 2025 Con- ference of the Nations of the Americas Cha...

  62. [70]

    Running Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang, Handi Chen, Weipeng Deng, Luyao Jin, Xiaojuan Qi, Xun Qian, and Edith CH Ngai. 2025. NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding.arXiv preprint arXiv:2508...

  63. [71]

    The Phenomenom of Light Reflection

    Hangyu Zhou, Yuichiro Fujimoto, Masayuki Kanbara, and Hirokazu Kato. 2021. Virtual reality as a reflection technique for public speaking training.Applied Sciences11, 9 (2021), 3988. UIST ’26, November 02–05, 2026, Detroit, MI, USA Fan et al. A Marzano’s Nine Events Table 2 sho...

  64. [72]

    Meilan Zhang, Mary Lundeberg, Matthew J Koehler, and Jan Eberhardt. 2011. Understanding affordances and challenges of three types of video for teacher professional development.Teaching and teacher education27, 2 (2011), 454–462

  65. [2014]

    IVE: an integrated interactive video-based learning environment

    L. IVE: an integrated interactive video-based learning environment. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 3399–3402

  66. [2018]

    Facilitating video analysis for teacher development: A systematic review of the research.Journal of Technology and Teacher Education26, 2 (2018), 185–216

  67. [2024]

    InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems

    ClassInSight: Designing Conversation Support Tools to Visualize Classroom Discussion for Personalized Teacher Professional Development. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–15

  68. [2025]

    InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems

    MindMate: Exploring the Effect of Conversational Agents on Reflective Writing. InProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–9

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.