REVIEW 4 major objections 6 minor 231 references
AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This dissertation claims that deep-learning systems at three learning stages—phoneme-level adaptive playback, multimodal video summarization, and unlabeled-data pronunciation feedback—can make audio-visual learning faster and more…
desk verdict A three-system dissertation with a genuinely novel AIxSpeed core, but the load-bearing evaluation of that core is partly circular and under-tested; it deserves review, not rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The object carrying the speed argument is the shared Wav2Vec2-based network in AIxSpeed, trained with Loss = Lossspeed + λLossctc: one head outputs a per-segment playback factor and the other performs CTC speech recognition, so a single model decides how fast each phoneme can go and verifies that the accelerated speech stays recognizable. The object carrying the summarization argument is FastPerson's video-to-video pipeline: chapter segmentation from scene and silence detection, visual metadata from OCR and object detection, transcription via Whisper, LLM-generated summaries, and VITS voice cloning to keep the narrator's voice continuous with the original. The object carrying the feedback argument is Profy's classifier, trained on unannotated learner and native speech, which scores utterances and visualizes the waveform regions the classifier attends to plus distances in latent space.
What would settle it
Take the same sentence materials from LibriSpeech and UME-ERJ, run AIxSpeed's per-phoneme speed profile, and have native and non-native listeners transcribe the output at word or phoneme level. If human accuracy does not track ASR confidence in those phoneme-level segments, or if the r≈0.998 correlation from the sentence-level pilot drops for non-native speech, then the claimed intelligibility preservation is unsupported. A second check: a quiz-based replication of FastPerson on unseen lectures with summary-only viewing would test whether the 53% time saving hides comprehension differences beyond the two videos used.
Extended reading notes
Core claim
The central claim is that speech-recognition confidence is a usable proxy for human listening difficulty at speeds above 1.0x, and that this proxy, a multimodal summarizer with voice preservation, and a self-supervised proficiency model each carry their respective stage of the learning cycle. The pilot study reports a correlation of r=0.9977 between human transcription accuracy and ASR accuracy across the tested speeds, and the full system is trained with a combined loss that maximizes speed while minimizing CTC recognition error, giving average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ with higher mean opinion scores than matched constant-speed playback. FastPerson is claimed to reduce viewing time by 53% on educational videos with quiz scores statistically indistinguishable from normal playback, and Profy is claimed to improve pronunciation intelligibility with pre- and post-practice confidence intervals that do not overlap the elicited-imitation baseline. The paper frames these three results as components of an integrated AI-Guided Learning loop rather than as standalone tools.
Load-bearing premise
The speed-optimization system stands or falls on the assumption that a speech recognizer's confidence remains a reliable proxy for human listening difficulty when applied at phoneme granularity to speech outside the tested English corpora, including non-native speech; the paper validates the proxy only at coarser granularity and for the specific datasets.
Editorial extensions
If this is right
- Learners using AIxSpeed would get roughly 23% shorter listening sessions at the same average quality, since 1.3x playback turns a 60-minute lecture into about 46 minutes.
- Video platforms could offer per-chapter summary versions that halve viewing time while quiz-measured comprehension stays statistically unchanged, with learners able to drill into full chapters when needed.
- Pronunciation practice with model-derived localized feedback can improve intelligibility without requiring transcribed error annotations, lowering the cost of building feedback systems for new languages or skills.
- The three systems compose a single learning loop: consume faster, understand via summaries or full segments, imitate with feedback, then return to relevant segments.
Reading between the lines
- A direct consequence the author leaves implicit is that AIxSpeed's speed profile could be personalized by raising or lowering the intelligibility threshold; the paper notes personalization as future work but does not test it.
- Because the pilot correlation was measured at utterance and sentence level, the strongest untested extension is the assumption that the same proxy holds at phoneme boundaries; a per-phoneme human transcription study would settle it.
- FastPerson's fixed summary-length weights (0.3 for audio, 2.5 for visual) suggest a testable extension: adapting those weights to learner proficiency or content type could yield further time savings or comprehension gains, which the dissertation does not test.
- The same Profy mechanism—a classifier trained on unannotated expert versus learner data with region highlighting—could transfer to other imitation domains such as music or movement, though the paper evaluates only English pronunciation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The dissertation proposes an AI-guided learning framework organized around three stages (Consume, Understand, Imitate) and instantiates it in three systems: AIxSpeed, which adapts audio playback speed at the phoneme level using speech-recognition confidence as a proxy for listening difficulty; FastPerson, which creates chapter-based video summaries preserving both visual and spoken information and uses voice cloning for continuity; and Profy, which provides pronunciation feedback from largely unannotated audio by highlighting classifier-emphasized waveform regions and latent-space distances. Reported evaluations include: AIxSpeed achieving average playback factors of 1.30x (LibriSpeech) and 1.29x (UME-ERJ) with higher mean opinion scores than matched constant-speed playback; FastPerson reducing viewing time by 53% with no statistically significant quiz-score difference; and Profy showing observed intelligibility improvements with non-overlapping pre/post confidence intervals.
Significance. If the central claims held, the work would make a useful contribution to human-computer interaction and educational technology: it demonstrates concrete, deployable prototypes that address real bottlenecks in audio-visual learning, and it integrates self-supervised speech representations, LLM-based multimodal summarization, and low-resource pronunciation feedback. The manuscript is unusually transparent: it repeatedly states where evidence is missing, e.g., that the AIxSpeed MOS differences were not significance-tested, that the UME-ERJ result does not establish human intelligibility, and that Profy's highlighted regions were not independently validated as phonetic error diagnoses. That transparency is a strength, but it also means the abstract and chapter-level summaries frequently assert more than the evidence supports. The framework itself (Consume-Understand-Imitate) is a reasonable organizing device, and each prototype addresses a genuine gap, so the work has clear value as a systems/HCI contribution if the claims are appropriately qualified.
major comments (4)
- [§3.5.1 and §3.3.2] The technical evaluation of AIxSpeed is partly self-referential. The playback speed adjuster is trained jointly with the speech recognizer through the composite loss Loss = Lossspeed + λLossctc (§3.3.2), and the CER/WER evaluation in §3.5.1 uses "the speech recognizer used in our method" to score the modified audio. The lower CER/WER relative to constant-speed playback therefore reflects, at least in part, that the speed profile was optimized to be recognizable to this exact evaluator. The pilot survey in §3.2 validates only aggregate, whole-utterance correlation at constant speeds; it does not validate per-phoneme confidence as a proxy for human intelligibility on non-native or time-stretched speech. The claim that intelligibility is preserved requires either an independent recognizer of a different architecture or a direct human transcription test at phoneme/segment granularity.
- [§3.5.2] The higher mean opinion scores for AIxSpeed over matched constant-speed playback are reported without a significance test; the text itself states that "The source paper did not report a significance test for these differences." A 0.5-point and 0.8-point difference on a five-point MOS scale is not interpretable without a test or confidence interval, especially with 50 participants and 40 sentences. Additionally, MOS is a measure of perceived quality, not comprehension, so this result cannot by itself support the abstract's implication that listening comprehension is maintained.
- [§4.4.3] The FastPerson quiz result is internally consistent in reporting no statistically significant difference, but the conclusion that comprehension was retained is stronger than the evidence warrants. The comparison is between 19 and 21 participants, and the quiz-score standard deviations are large (e.g., Video 2: 0.67 ± 0.46 for FastPerson vs. 0.63 ± 0.39 for control). A non-significant t-test with these sample sizes does not establish equivalence; a non-inferiority test or a confidence interval for the quiz-score difference is needed. The viewing-time reduction (53%) is well supported by the reported t-tests, so the efficiency claim is robust; the claim of preserved understanding is not established at the same evidentiary level.
- [§5.4 and §5.5.2] The Profy evaluation involves only 10 learners and 5 raters, and the headline finding is presented as non-overlapping pre- and post-practice confidence intervals without reporting the interval values or a test statistic. The observed intelligibility gain could reflect practice effects or rater drift, and the elicited-imitation baseline is not accompanied by effect sizes. Furthermore, as the manuscript itself acknowledges in §2.8.2 and §5.5.2, the assumption that classifier-emphasized waveform regions and latent distances correspond to correctable pronunciation deviations was not independently validated. The chapter's central claim of improved pronunciation intelligibility is therefore plausible but not yet firmly supported; reporting raw scores, CIs, and an effect size would help, as would a control condition for the feedback mechanism itself.
minor comments (6)
- [§3.2] The formulas for WER and CER state the denominator as "number of correct words" (and characters); standard definitions divide by the total number of reference words/characters. As written, the metric is non-standard and should be clarified or corrected.
- [Table 3.1] The CER/WER comparisons in Table 3.1 report point estimates only; confidence intervals or significance tests for the differences would improve interpretability.
- [Abstract and §1.7.2] The abstract and the contribution list state that AIxSpeed "received higher mean opinion scores" and that UME-ERJ "suggests" improved listenability for non-native speech. Given the explicit caveats in §3.5.2 and §3.6.4, these statements should be qualified as observed differences without significance testing.
- [§4.4.3 and Figure 4.3] The text refers to the "right side" and "left side" of Figure 4.3, but the figure appears as a single panel of per-question correct-answer rates. Please align the description with the actual figure layout.
- [§4.2.4] The summary-length weights ws = 0.3 and wv = 2.5 are described as determined by preliminary experiments; reporting the criterion used (e.g., which pilot data, what target) would improve reproducibility.
- [§5.4] For the Profy result, please report the actual confidence interval values rather than only stating that they do not overlap, and specify the statistical procedure used to construct them.
Circularity Check
AIxSpeed's ASR-intelligibility result is circular: the speed adjuster is trained against the same recognizer that is then used to evaluate it, leaving only an un-significance-tested MOS comparison as external support.
-
fitted input called prediction
[Section 3.3.2 (Speech Recognizer), Section 3.4.1 (Implementation Details), Section 3.5.1 (Technical Evaluation)]
"In summary, the entire model is trained to minimize the following error function, which is a combination of this error function and the error function of the playback speed adjuster: Loss = Lossspeed + λLossctc. ... A standard Wav2Vec2-based speech recognition model, which was the speech recognizer used in our method, was used to compute the CER and WER for comparison. ... This indicates that it was more recognizable to the ASR model used for evaluation; it does not directly establish improved intelligibility or comprehension for human listeners, which requires a dedicated human study."
The playback-speed adjuster is trained with a joint objective that includes the recognizer's CTC loss (Loss = Lossspeed + λLossctc), and the recognizer's weights are held fixed while the adjuster is optimized against that recognizer. The technical evaluation then measures CER/WER with 'the speech recognizer used in our method'—the exact model whose loss was part of the training objective. Lower CER/WER for AIxSpeed than for matched constant-speed playback therefore partly reflects that the adjuster was fitted to make this recognizer succeed, not an independent measure of human intelligibility.
full rationale
The clearest circularity is localized to AIxSpeed. The speed adjuster is optimized with a loss that explicitly includes the speech recognizer's CTC loss, and the technical evaluation uses that same recognizer to compute CER/WER; the observed improvement over matched constant-speed playback is thus partly a self-consistency check of the training objective rather than an independent intelligibility result. The paper is transparent about this limitation in the quoted passages, but the chapter conclusion still leans on the ASR numbers to claim intelligibility preservation. The pilot survey correlating human and ASR transcription accuracy (r=0.9977) is independent external evidence, though it validates only aggregate speed conditions and not per-phoneme decisions. FastPerson's viewing-time reduction is a measured outcome rather than a fitted prediction, and its quiz scores are independent human outcomes; Profy is evaluated by external human raters. There is no load-bearing self-citation chain. Overall, one central technical result reduces by construction to its training objective, warranting a partial-circularity score of 6.
Assumptions & free parameters
free parameters (4)
- AIxSpeed loss weight lambda =
1e-7
- FastPerson summary length weights ws and wv =
ws=0.3, wv=2.5
- FastPerson segmentation thresholds =
Bhattacharyya 0.6; silence RMS below 5% for >1s
- FastPerson minimum summary length Nmin =
50 words
assumptions (6)
- domain assumption ASR confidence is a valid proxy for human listening difficulty for the purpose of setting playback speed.
- domain assumption The aggregate human-ASR correlation extrapolates to fine-grained per-phoneme speed decisions.
- domain assumption LLM-based summaries from transcripts, OCR, and object labels retain the content needed for quiz performance.
- domain assumption Voice cloning preserves sufficient speaker identity and prosody for seamless learning continuity.
- ad hoc to paper Classifier-emphasized waveform regions and latent distances correspond to correctable pronunciation deviations.
- domain assumption Absence of a statistically significant quiz-score difference in a small sample is treated as evidence that summarized viewing does not harm learning.
Cite this review
Pith. "Pith review of AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques." pith.science (2026). https://pith.science/paper/56IP7E6O
@misc{pith2026260808990,
author = {Pith},
title = {Pith review of: AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques},
year = {2026},
howpublished = {\url{https://pith.science/paper/56IP7E6O}},
note = {Machine review of arXiv:2608.08990}
}
read the original abstract
Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.
Figures
Figures from the paper (23 more)
Reference graph
Works this paper leans on
-
[1]
The effect of settings, educational level and tools on computer-assisted pronun- ciation training: A meta-analysis
Asma Almusharraf, Hassan Saleh Mahdi, Haifa Al-Nofaie, and Amal Aljasser. The effect of settings, educational level and tools on computer-assisted pronun- ciation training: A meta-analysis. Journal of Computer Assisted Learning , 40(4):1605–1615, 2024
2024
-
[2]
Advances in computer-based education: The Plato program will provide a major test of the educational and economic fea- sibility of this medium
Daniel Alpert and Donald Lester Bitzer. Advances in computer-based education: The Plato program will provide a major test of the educational and economic fea- sibility of this medium. Science, 167(3925):1582–1590, 1970
1970
-
[3]
Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Guidelines for human-AI interac- tion. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2019
2019
-
[4]
Computer-assisted pronunciation training: A systematic review
Moustafa Amrate and Pi-hua Tsai. Computer-assisted pronunciation training: A systematic review. ReCALL, 37(1):22–42, 2025. Published online 19 September 2024
2025
-
[5]
Anderson and Brian J
John R. Anderson and Brian J. Reiser. The LISP tutor. BYTE, 10(4):159–175, 1985
1985
-
[6]
Anderson and David R
Lorin W. Anderson and David R. Krathwohl, editors. A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objec- tives. Longman, New Y ork, 2001
2001
-
[7]
Metsai, Vasileios Mezaris, and Ioannis Patras
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, and Ioannis Patras. Video summarization using deep neural networks: A survey. Proceedings of the IEEE , 109(11):1838–1863, 2021
2021
-
[8]
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis T yers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4218–4222. European Language Resources Association, 2020
2020
Show all 231 references
-
[9]
Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Y ongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi
Sercan Ö. Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Y ongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi. Deep voice: Real-time neural text- to-speech. In Proceedings of the International Conference...
2017
-
[10]
ViViT: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. ViViT: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 6816–6826, 2021. 90
2021
-
[11]
SpeechSkimmer: A system for interactively skimming recorded speech
Barry Arons. SpeechSkimmer: A system for interactively skimming recorded speech. ACM Transactions on Computer-Human Interaction, 4(1):3–38, 1997
1997
-
[12]
Cromley, and Diane Seibert
Roger Azevedo, Jennifer G. Cromley, and Diane Seibert. Does adaptive scaffold- ing facilitate students’ ability to regulate their learning with hypermedia? Con- temporary Educational Psychology, 29(3):344–370, 2004
2004
-
[13]
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, volume 33, pages 12449– 12460. Curran Associates, Inc., 2020
2020
-
[14]
Neural machine transla- tion by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Y oshua Bengio. Neural machine transla- tion by jointly learning to align and translate. In Proceedings of the International Conference on Learning Representations, 2015
2015
-
[15]
Speaker recognition based on deep learning: An overview
Zhongxin Bai and Xiao-Lei Zhang. Speaker recognition based on deep learning: An overview. Neural Networks, 140:65–99, 2021
2021
-
[16]
Social learning theory
Albert Bandura. Social learning theory . Prentice Hall, 1977
1977
-
[17]
Self-efficacy: The exercise of control
Albert Bandura. Self-efficacy: The exercise of control . W. H. Freeman, 1997
1997
-
[18]
Meeting the universe halfway: Quantum physics and the entangle- ment of matter and meaning
Karen Barad. Meeting the universe halfway: Quantum physics and the entangle- ment of matter and meaning . Duke University Press, 2007
2007
-
[19]
Text summarization using large language mod- els: A comparative study of MPT-7b-instruct, Falcon-7b-instruct, and OpenAI Chat-GPT models
Lochan Basyal and Mihir Sanghvi. Text summarization using large language mod- els: A comparative study of MPT-7b-instruct, Falcon-7b-instruct, and OpenAI Chat-GPT models. arXiv, 2023
2023
-
[20]
Enhancing learning ex- periences: EEG-based passive BCI system adapts learning speed to cognitive load in real-time, with motivation as catalyst
Noémie Beauchemin, Patrick Charland, Alexander Karran, Jared Boasen, Bella Tadson, Sylvain Sénécal, and Pierre-Majorique Léger. Enhancing learning ex- periences: EEG-based passive BCI system adapts learning speed to cognitive load in real-time, with motivation as catalyst. Fro...
2024
-
[21]
Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the ACM Conference on Fairness, Accountability, and Trans- parency, pages 610–623. Association for Computing M...
2021
-
[22]
Abstractive video lecture summarization: Applications and future prospects
Irene Benedetto, Moreno La Quatra, Luca Cagliero, Lorenzo Canale, and Laura Farinetti. Abstractive video lecture summarization: Applications and future prospects. Education & Information Technologies, 29(3):2951–2971, 2024
2024
-
[23]
The evidence for ’flipping out’: A systematic review of the flipped classroom in nursing education
Vasiliki Betihavas, Heather Bridgman, Rachel Kornhaber, and Merylin Cross. The evidence for ’flipping out’: A systematic review of the flipped classroom in nursing education. Nurse Education Today, 38:15–21, 2016
2016
-
[24]
Bitzer, Peter G
Donald L. Bitzer, Peter G. Braunfeld, and Wayne W. Lichtenberger. PLATO: An automatic teaching device. IRE Transactions on Education, 4(4):157–161, 1961
1961
-
[25]
Duolingo language report 2024, 2024
Cindy Blanco. Duolingo language report 2024, 2024. Retrieved May 25, 2025, from https://blog.duolingo.com/2024-duolingo-language-report/
2024
-
[26]
Benjamin S. Bloom. The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6):4– 16, 1984. 91
1984
-
[27]
Outline of a theory of practice
Pierre Bourdieu. Outline of a theory of practice . Cambridge University Press, 1977
1977
-
[28]
Cynthia J. Brame. Effective educational videos: Principles and guidelines for maximizing student learning from video content. CBELife Sciences Education, 15(4):es6, 2016
2016
-
[29]
PTeacher: A computer-aided personalized pronunciation training system with exaggerated audio-visual corrective feedback
Y aohua Bu, Tianyi Ma, Weijun Li, Hang Zhou, Jia Jia, Shengqi Chen, Kaiyuan Xu, Dachuan Shi, Haozhe Wu, Zhihan Y ang, Kun Li, Zhiyong Wu, Yuanchun Shi, Xi- aobo Lu, and Ziwei Liu. PTeacher: A computer-aided personalized pronunciation training system with exaggerated audio-visu...
2021
-
[30]
XTTS: A massively multilingual zero-shot text-to-speech model
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Lo- gan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, and Ju- lian Weber. XTTS: A massively multilingual zero-shot text-to-speech model. In Proceedings of INTERSPEECH 2024 , pages 4978–...
2024
-
[31]
Seamful and seamless design in ubiquitous computing
Matthew Chalmers and Ian MacColl. Seamful and seamless design in ubiquitous computing. In Workshop on At the Crossroads: The Interaction of HCI and Sys- tems Issues in UbiComp (UbiComp 2003 Workshop) , 2003. Position paper
2003
-
[32]
Kumar, Rhea Varkhedi, and Dillon H
Ashley Chen, Suchita E. Kumar, Rhea Varkhedi, and Dillon H. Murphy. The effect of playback speed and distractions on the comprehension of audio and audio-visual materials. Educational Psychology Review, 36:79, 2024
2024
-
[33]
Video browse - a study of user behavior in online VoD services
Liang Chen, Yipeng Zhou, and Dah Ming Chiu. Video browse - a study of user behavior in online VoD services. In Proceedings of the International Conference on Computer Communication and Networks, pages 1–7, 2013
2013
-
[34]
Computer assisted pronunciation training (CAPT): A systematic review of studies from 2012 to 2021
Xu Chen, Jie Mu, and Tingting Zhang. Computer assisted pronunciation training (CAPT): A systematic review of studies from 2012 to 2021. In Proceedings of the International Conference on Computers in Education , pages 575–580. Asia- Pacific Society for Computers in Education, 2022
2012
-
[35]
MultiPA: A multi-task speech pronunciation assessment model for open response scenarios
Yu-Wen Chen, Zhou Yu, and Julia Hirschberg. MultiPA: A multi-task speech pronunciation assessment model for open response scenarios. In Proceedings of INTERSPEECH, pages 297–301, 2024
2024
-
[36]
The ICAP framework: Linking cognitive engagement to active learning outcomes
Michelene TH Chi and Ruth Wylie. The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4):219– 243, 2014
2014
-
[37]
MixT: Automatic generation of step-by-step mixed media tutorials
Pei- Yu Chi, Sally Ahn, Amanda Ren, Mira Dontcheva, Wilmot Li, and Björn Hart- mann. MixT: Automatic generation of step-by-step mixed media tutorials. In Proceedings of the Annual ACM Symposium on User Interface Software and Tech- nology, pages 93–102, 2012
2012
-
[38]
VIVID: Human-AI col- laborative authoring of vicarious dialogues from lecture videos
Seulgi Choi, Hyewon Lee, Y oonjoo Lee, and Juho Kim. VIVID: Human-AI col- laborative authoring of vicarious dialogues from lecture videos. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–26. As- sociation for Computing Machinery, 2024
2024
-
[39]
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, KyungHyun Cho, and Y oshua Bengio. Attention-based models for speech recognition. In Proceedings 92 of the International Conference on Neural Information Processing Systems, pages 577–585, 2015
2015
-
[40]
The extended mind
Andy Clark and David Chalmers. The extended mind. Analysis, 58(1):7–19, 1998
1998
-
[41]
Reac- tive video: Adaptive video playback based on user motion for supporting physical activity
Christopher Clarke, Larissa Pschetz, Mor Trope, and Dave Murray-Rust. Reac- tive video: Adaptive video playback based on user motion for supporting physical activity. In Proceedings of the International Conference on Intelligent User Inter- faces, pages 196–208, 2020
2020
-
[42]
Measuring mind wandering during online lectures assessed with EEG
Colin Conrad and Aaron Newman. Measuring mind wandering during online lectures assessed with EEG. Frontiers in Human Neuroscience, 15:697532, 2021
2021
-
[43]
The effects of time-compressed speech on native and EFL listening comprehension
Linda Conrad. The effects of time-compressed speech on native and EFL listening comprehension. Studies in Second Language Acquisition , 11(1):1–16, 1989
1989
-
[44]
Fergus I. M. Craik and Robert S. Lockhart. Levels of processing: A framework for memory research. Journal of Verbal Learning and Verbal Behavior , 11(6):671– 684, 1972
1972
-
[45]
Artificial intelligence in higher education: The state of the field
Helen Crompton and Diane Burke. Artificial intelligence in higher education: The state of the field. International Journal of Educational Technology in Higher Education, 20(1):22, 2023
2023
-
[46]
The friendly orange glow: The untold story of the PLATO system and the dawn of cyberculture
Brian Dear. The friendly orange glow: The untold story of the PLATO system and the dawn of cyberculture. Pantheon Books, 2017
2017
-
[47]
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language T...
2019
-
[48]
Dillahunt, Zengguang Wang, and Stephanie D
Tawanna R. Dillahunt, Zengguang Wang, and Stephanie D. Teasley. Democra- tizing higher education: Exploring MOOC use among those who cannot afford a formal education. The International Review of Research in Open and Distributed Learning, 15(5):177–196, 2014
2014
-
[49]
Video browse by direct manipulation
Pierre Dragicevic, Gonzalo Ramos, Jacobo Bibliowitcz, Derek Nowrouzezahrai, Ravin Balakrishnan, and Karan Singh. Video browse by direct manipulation. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 237–246, 2008
2008
-
[50]
Drucker, Asta Glatzer, Steven De Mar, and Curtis Wong
Steven M. Drucker, Asta Glatzer, Steven De Mar, and Curtis Wong. SmartSkip: Consumer level browse and skipping of digital video content. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 219–226, 2002
2002
-
[51]
Why college students watch streaming drama at higher playback speed: The uses and gratifications perspective
Songshuang Duan and Xiaoqian Chen. Why college students watch streaming drama at higher playback speed: The uses and gratifications perspective. In Pro- ceedings of the International Joint Conference on Information, Media and Engi- neering, 2019
2019
-
[52]
Anders Ericsson, Ralf T
K. Anders Ericsson, Ralf T. Krampe, and Clemens Tesch-Römer. The role of deliberate practice in the acquisition of expert performance.Psychological Review, 100(3):363–406, 1993. 93
1993
-
[53]
Metacognition and self-regulation
Evidence for Learning. Metacognition and self-regulation. Teaching and Learning Toolkit, 2021. Review last updated July 2021
2021
-
[54]
Investigating the effects of artificial intelligence-assisted language learning strategies on cognitive load and learning outcomes: A comparative study
Lijuan Feng. Investigating the effects of artificial intelligence-assisted language learning strategies on cognitive load and learning outcomes: A comparative study. Journal of Educational Computing Research , 62(8):1741–1774, 2025. First pub- lished online 31 August 2024; iss...
2025
-
[55]
Effects of experience on non-native speakers’ production and perception of English vowels
James Emil Flege, Ocke-Schwen Bohn, and Sunyoung Jang. Effects of experience on non-native speakers’ production and perception of English vowels. Journal of Phonetics, 25(4):437–470, 1997
1997
-
[56]
Au- tomatic speech recognition predicts speech intelligibility and comprehension for listeners with simulated age-related hearing loss
Lionel Fontan, Isabelle Ferrané, Jérôme Farinas, Julien Pinquier, Julien Tardieu, Cynthia Magnen, Pascal Gaillard, Xavier Aumont, and Christian Füllgrabe. Au- tomatic speech recognition predicts speech intelligibility and comprehension for listeners with simulated age-related ...
2017
-
[57]
Adam Fouse, Nadir Weibel, Edwin Hutchins, and James D. Hollan. ChronoViz: A system for supporting navigation of time-coded data. In Proceedings of the CHI Conference on Human Factors in Computing Systems Extended Abstracts , pages 299–304, 2011
2011
-
[58]
Pedagogy of the oppressed
Paulo Freire. Pedagogy of the oppressed. Seabury Press, 1970
1970
-
[59]
The perceptual learning of time- compressed speech: A comparison of training protocols with different levels of difficulty
Y afit Gabay, Avi Karni, and Karen Banai. The perceptual learning of time- compressed speech: A comparison of training protocols with different levels of difficulty. PLoS ONE, 12(5):e0176488, 2017
2017
-
[60]
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren. DARPA TIMIT acoustic phonetic continuous speech corpus CDROM. Technical Report NISTIR-4930, National Institute of Standards and Technology (NIST), 1993
1993
-
[61]
Giegerich
Heinz J. Giegerich. English phonology: An introduction . Cambridge Textbooks in Linguistics. Cambridge University Press, 1992
1992
-
[62]
The complete HyperCard handbook
Danny Goodman. The complete HyperCard handbook. Bantam Books, 1987
1987
-
[63]
Graves, S
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural nets. In Proceedings of the International Conference on Machine Learning , 2006
2006
-
[64]
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In Proceedings of the International Confer- ence on Acoustics, Speech and Signal Processing , pages 6645–6649, 2013
2013
-
[65]
A survey on self-supervised learning: Algorithms, applications, and fu- ture trends
Jie Gui, Tuo Chen, Jing Zhang, Qiong Cao, Zhenan Sun, Hao Luo, and Dacheng Tao. A survey on self-supervised learning: Algorithms, applications, and fu- ture trends. IEEE Transactions on Pattern Analysis and Machine Intelligence , 46(12):9052–9071, 2024
2024
-
[66]
Guo, Juho Kim, and Rob Rubin
Philip J. Guo, Juho Kim, and Rob Rubin. How video production affects student engagement: An empirical study of MOOC videos. In Proceedings of the ACM Conference on Learning at Scale , pages 41–50. ACM, 2014. 94
2014
-
[67]
M. P . J. Habgood and S. E. Ainsworth. Motivating children to learn effectively: Exploring the value of intrinsic integration in educational games. The Journal of the Learning Sciences , 20(2):169–206, 2011
2011
-
[68]
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and An- drew Y . Ng. Deep speech: Scaling up end-to-end speech recognition. arXiv, 2014
2014
-
[69]
Exploring collaborative decision- making: A quasi-experimental study of human and generative AI interaction.Tech- nology in Society, 78:102662, 2024
Xinyue Hao, Emrah Demir, and Daniel Eyers. Exploring collaborative decision- making: A quasi-experimental study of human and generative AI interaction.Tech- nology in Society, 78:102662, 2024
2024
-
[70]
Hardison
Debra M. Hardison. Generalization of computer-assisted prosody training: Quan- titative and qualitative findings. Language Learning & Technology , 8(1):34–52, 2004
2004
-
[71]
Hartshorne, Joshua B
Joshua K. Hartshorne, Joshua B. Tenenbaum, and Steven Pinker. A critical pe- riod for second language acquisition: Evidence from 2/3 million English speakers. Cognition, 177:263–277, 2018
2018
-
[72]
Teaching machines to read and com- prehend
Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and com- prehend. In Proceedings of the International Conference on Neural Information Processing Systems, pages 1693–1701, 2015
2015
-
[73]
Students’ and instructors’ use of mas- sive open online courses (MOOCs): Motivations and challenges
Khe Foon Hew and Wing Sum Cheung. Students’ and instructors’ use of mas- sive open online courses (MOOCs): Motivations and challenges. Educational Research Review, 12:45–58, 2014
2014
-
[74]
EgoScanning: Quickly scanning first-person videos with egocentric elastic timelines
Keita Higuchi, Ryo Y onetani, and Y oichi Sato. EgoScanning: Quickly scanning first-person videos with egocentric elastic timelines. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 6536–6546. Associ- ation for Computing Machinery, 2017
2017
-
[75]
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon. Deep metric learning using triplet network. Similarity- Based Pattern Recognition, pages 84–92, 2015
2015
-
[76]
Designing for human-AI complementarity in K-12 education
Kenneth Holstein and Vincent Aleven. Designing for human-AI complementarity in K-12 education. AI Magazine, 43(2):239–248, 2022
2022
-
[77]
The struggle for recognition: The moral grammar of social con- flicts
Axel Honneth. The struggle for recognition: The moral grammar of social con- flicts. MIT Press, 1995
1995
-
[78]
Hershey, Tim K
Chiori Hori, Takaaki Hori, Teng- Y ok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In Proceedings of the IEEE International Conference on Computer Vision, pages 4193–4202, 2017
2017
-
[79]
Principles of mixed-initiative user interfaces
Eric Horvitz. Principles of mixed-initiative user interfaces. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 159–166, 1999
1999
-
[80]
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Y ao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Trans- actions on Audio, Speech and Language Processing , ...
2021
-
[81]
Cognition in the Wild
Edwin Hutchins. Cognition in the Wild . MIT Press, 1995. 95
1995
-
[82]
Deschooling society
Ivan Illich. Deschooling society. Harper & Row, 1971
1971
-
[83]
Designing participatory AI: Cre- ative professionals`worries and expectations about generative AI
Nanna Inie, Jeanette Falk, and Steve Tanimoto. Designing participatory AI: Cre- ative professionals`worries and expectations about generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems Extended Ab- stracts, pages 1–8, 2023
2023
-
[84]
Tangible bits: Towards seamless interfaces be- tween people, bits and atoms
Hiroshi Ishii and Brygg Ullmer. Tangible bits: Towards seamless interfaces be- tween people, bits and atoms. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 234–241, 1997
1997
-
[85]
Word perception in fast speech: Artificially time-compressed vs
Esther Janse. Word perception in fast speech: Artificially time-compressed vs. naturally produced fast speech. Speech Communication, 42:155–173, 2004
2004
-
[86]
Speech recognition performance as an ef- fective perceived quality predictor
Wenyu Jiang and Henning Schulzrinne. Speech recognition performance as an ef- fective perceived quality predictor. InProceedings of the Tenth IEEE International Workshop on Quality of Service , pages 269–275. IEEE, 2002
2002
-
[87]
Us- ing ChatGPT for course curriculum design: A systematic review
Michelle Celine J”orgens, Florian Beier, Sebastian Kreibich, and Dirk Werth. Us- ing ChatGPT for course curriculum design: A systematic review. In The Paris Conference on Education 2024: Official Conference Proceedings, pages 549–561, 2024
2024
-
[88]
A survey of speaker recognition: Fundamental theories, recognition methods and opportunities
Muhammad Mohsin Kabir, Muhammad Firoz Mridha, Jungpil Shin, Israt Jahan, and Abu Quwsar Ohi. A survey of speaker recognition: Fundamental theories, recognition methods and opportunities. IEEE Access, 9:79236–79263, 2021
2021
-
[89]
The expertise rever- sal effect
Slava Kalyuga, Paul Ayres, Paul Chandler, and John Sweller. The expertise rever- sal effect. Educational Psychologist, 38(1):23–31, 2003
2003
-
[90]
Educational scalability in MOOCs: Analysing instructional designs to find best practices
Julia Kasch, Peter Van Rosmalen, and Marco Kalz. Educational scalability in MOOCs: Analysing instructional designs to find best practices. Computers & Education, 161:104054, 2021
2021
-
[91]
ChatGPT for good? on opportunities and challenges of large language models for education
Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Saile...
2023
-
[92]
Efficient video viewing system for racquet sports with automatic summarization focusing on rally scenes
Shunya Kawamura, Tsukasa Fukusato, Tatsunori Hirai, and Shigeo Morishima. Efficient video viewing system for racquet sports with automatic summarization focusing on rally scenes. In Proceedings of the ACM SIGGRAPH Posters, 2014
2014
-
[93]
Robin H. Kay. Exploring the use of video podcasts in education: A comprehensive review of the literature. Computers in Human Behavior , 28(3):820–831, 2012
2012
-
[94]
Explainable artificial intelligence in education: A com- prehensive review
Hassan Khosravi, Simon Buckingham Shum, Guanliang Chen, Cristina Conati, Yi- Shan Tsai, Judy Kay, Simon Knight, Roberto Martinez-Maldonado, Shazia Sadiq, and Dragan Gašević. Explainable artificial intelligence in education: A com- prehensive review. Computers and Education: Ar...
2022
-
[95]
Digital game-based learning: Towards an experiential gaming model
Kristian Kiili. Digital game-based learning: Towards an experiential gaming model. The Internet and Higher Education , 8(1):13–24, 2005. 96
2005
-
[96]
Automatic pronunciation assessment using self-supervised speech representation learning
Eesung Kim, Jae-Jin Jeon, Hyeji Seo, and Hoon Kim. Automatic pronunciation assessment using self-supervised speech representation learning. In Proceedings of INTERSPEECH, pages 1411–1415. ISCA, 2022
2022
-
[97]
Generic speech summarization of transcribed lecture videos: Using tags and their semantic relations
Hyun Hee Kim and Y ong Ho Kim. Generic speech summarization of transcribed lecture videos: Using tags and their semantic relations. Journal of the Association for Information Science and Technology, 67(2):366–379, 2016
2016
-
[98]
Guo, Daniel T
Juho Kim, Philip J. Guo, Daniel T. Seaton, Piotr Mitros, Krzysztof Z. Gajos, and Robert C. Miller. Understanding in-video dropouts and interaction peaks in online lecture videos. In Proceedings of the First ACM Conference on Learning at Scale, pages 31–40. ACM, 2014
2014
-
[99]
Kingma and Jimmy Ba
Diederik P . Kingma and Jimmy Ba. Adam: A method for stochastic optimiza- tion. In Proceedings of the International Conference on Learning Representations, 2015
2015
-
[100]
David A. Kolb. Experiential Learning: Experience as the Source of Learning and Development. Prentice-Hall, Englewood Cliffs, NJ, 1984
1984
-
[101]
Kulik and J
James A. Kulik and J. D. Fletcher. Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1):42–78, 2016
2016
-
[102]
CinemaGazer: A system for watching video at very high speed
Kazutaka Kurihara. CinemaGazer: A system for watching video at very high speed. In Proceedings of the Workshop on Advanced Visual Interfaces , 2012
2012
-
[103]
Is faster better? a study of video playback speed
David Lang, Guanliang Chen, Kathy Mirzaei, and Andreas Paepcke. Is faster better? a study of video playback speed. In Proceedings of the International Conference on Learning Analytics & Knowledge, pages 260–269. Association for Computing Machinery, 2020
2020
-
[104]
Situated learning: Legitimate peripheral partici- pation
Jean Lave and Etienne Wenger. Situated learning: Legitimate peripheral partici- pation. Cambridge University Press, 1991
1991
-
[105]
Berg, and Mohit Bansal
Jie Lei, Tamara L. Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. InAdvances in Neural Information Processing Systems, volume 34, pages 11846–11858, 2021
2021
-
[106]
Chen, and Chin-Hui Lee
Wei Li, Kehuang Li, Sabato Marco Siniscalchi, Nancy F. Chen, and Chin-Hui Lee. Detecting mispronunciations of L2 learners and providing corrective feed- back using knowledge-guided and data-driven decision trees. In Proceedings of INTERSPEECH, 2016
2016
-
[107]
Enhancing length generalization for attention based knowledge tracing models with linear biases
Xueyi Li, Y ouheng Bai, Teng Guo, Zitao Liu, Y aying Huang, Xiangyu Zhao, Feng Xia, Weiqi Luo, and Jian Weng. Enhancing length generalization for attention based knowledge tracing models with linear biases. In Proceedings of the Thirty- Third International Joint Conference on ...
2024
-
[108]
Focal loss for dense object detection
Tsung- Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, pages 2980–2988, 2017. 97
2017
-
[109]
Audio self-supervised learn- ing: A survey
Shuo Liu, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Kun Qian, Xin Jing, Alexander Kathan, Bin Hu, and Bjoern W Schuller. Audio self-supervised learn- ing: A survey. Patterns, 3(12):100616, 2022
2022
-
[110]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Pro- ceedings of the 7th International Conference on Learning Representations , 2019
2019
-
[111]
The modality principle in multimedia learning
Renae Low and John Sweller. The modality principle in multimedia learning. In Richard E. Mayer, editor,The Cambridge Handbook of Multimedia Learning (2nd ed.), pages 227–246. Cambridge University Press, 2014
2014
-
[112]
Automatic speech recognition: A survey
Mishaim Malik, Muhammad Malik, Khawar Mehmood, and Imran Makhdoom. Automatic speech recognition: A survey. Multimedia Tools and Applications , 80:9411–9457, 2021
2021
-
[113]
Florence Martin, Ting Sun, and Carl D. Westine. A systematic review of research on online teaching and learning from 2009 to 2018. Computers & Education , 159:104009, 2020
2009
-
[114]
Individual differences in working memory capacity moderate effects of post-learning activity on memory consolidation over the long term
Markus Martini, Robert Marhenke, Caroline Martini, Sonja Rossi, and Pierre Sachse. Individual differences in working memory capacity moderate effects of post-learning activity on memory consolidation over the long term. Scientific Re- ports, 10(1):17976, 2020
2020
-
[115]
Swift: Reducing the effects of latency in online video scrubbing
Justin Matejka, Tovi Grossman, and George Fitzmaurice. Swift: Reducing the effects of latency in online video scrubbing. In Proceedings of the CHI Confer- ence on Human Factors in Computing Systems , pages 637–646. Association for Computing Machinery, 2012
2012
-
[116]
Mayer, editor
Richard E. Mayer, editor. The Cambridge handbook of multimedia learning. Cam- bridge University Press, 2005
2005
-
[117]
Richard E. Mayer. Cognitive theory of multimedia learning. In Richard E. Mayer, editor, The Cambridge handbook of multimedia learning, pages 31–48. Cambridge University Press, 2005
2005
-
[118]
Richard E. Mayer. Multimedia learning . Cambridge University Press, 3rd ed. edition, 2020
2020
-
[119]
Mayer, Kristina Sobko, and Patricia D
Richard E. Mayer, Kristina Sobko, and Patricia D. Mautone. Social cues in mul- timedia learning: Role of speaker’s voice. Journal of Educational Psychology , 95(2):419–425, 2003
2003
-
[120]
Montreal Forced Aligner: Trainable text-speech alignment us- ing Kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Mor- gan Sonderegger. Montreal Forced Aligner: Trainable text-speech alignment us- ing Kaldi. In Proceedings of INTERSPEECH, pages 498–502, 2017
2017
-
[121]
Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto
Brian McFee, Colin Raffel, Dawen Liang, Daniel P . Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. Librosa: Audio and music signal analysis in Python. In Proceedings of the Python in Science Conference , pages 18–24, 2015
2015
-
[122]
Development of En- glish speech database spoken by Japanese learners
Nobuaki Minematsu, Y oshihiro Tomiyama, Kei Y oshimoto, Katsumasa Shimizu, Seiichi Nakagawa, Masatake Dantsuji, and Shozo Makino. Development of En- glish speech database spoken by Japanese learners. In Proceedings of the CO- COSDA Workshop 2001, pages 76–81, 2001. 98
2001
-
[123]
English speech database read by Japanese learners for CALL system development
Nobuaki Minematsu, Y oshihiro Tomiyama, Kei Y oshimoto, Katsumasa Shimizu, Seiichi Nakagawa, Masatake Dantsuji, and Shozo Makino. English speech database read by Japanese learners for CALL system development. InProceedings of the International Conference on Language Resources ...
2002
-
[124]
Punya Mishra and Matthew J. Koehler. Technological pedagogical content knowledge: A framework for teacher knowledge. Teachers College Record , 108(6):1017–1054, 2006
2006
-
[125]
Roxana Moreno and Richard E. Mayer. Cognitive principles of multimedia learn- ing: The role of modality and contiguity. Journal of Educational Psychology , 91(2):358–368, 1999
1999
-
[126]
Multimodal data fusion in learning analyt- ics: A systematic review
Su Mu, Meng Cui, and Xiaodi Huang. Multimodal data fusion in learning analyt- ics: A systematic review. Sensors, 20(23):6856, 2020
2020
-
[127]
Murphy, Kara M
Dillon H. Murphy, Kara M. Hoover, Karina Agadzhanyan, Jesse C. Kuehn, and Alan D. Castel. Learning in double time: The effect of lecture video speed on immediate and delayed comprehension. Applied Cognitive Psychology, 36(1):69– 82, 2022
2022
-
[128]
The pedagogy– technology interface in computer assisted pronunciation training
Ambra Neri, Catia Cucchiarini, Helmer Strik, and Lou Boves. The pedagogy– technology interface in computer assisted pronunciation training. Computer As- sisted Language Learning, 15(5):441–467, 2002
2002
-
[129]
The effective- ness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis
Thuy Thi-Nhu Ngo, Howard Hao-Jan Chen, and Kyle Kuo-Wei Lai. The effective- ness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis. ReCALL, 36(1):4–21, 2024
2024
-
[130]
Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran
Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. Measuring calibration in deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages 38–41, 2019
2019
-
[131]
Donald A. Norman. The design of everyday things: Revised and expanded edition . Basic Books, 2013
2013
-
[132]
GPT-4V(ision) system card
OpenAI. GPT-4V(ision) system card. Technical report, OpenAI, 2023
2023
-
[133]
Pacheco, Charley W
Matheus M. Pacheco, Charley W. Lafe, and Karl M. Newell. Search strategies in the perceptual-motor workspace and the acquisition of coordination, control, and skill. Frontiers in Psychology, 10:1874, 2019
2019
-
[134]
The vowel game: Continuous real- time visualization for pronunciation learning with vowel charts
Annu Paganus, Vesa-Petteri Mikkonen, Tomi Mäntylä, Sami Nuuttila, Jouni Isoaho, Olli Aaltonen, and Tapio Salakoski. The vowel game: Continuous real- time visualization for pronunciation learning with vowel charts. In Advances in Natural Language Processing, pages 696–703. Spri...
2006
-
[135]
Mul- timodal abstractive summarization for How2 videos
Shruti Palaskar, Jindřich Libovický, Spandana Gella, and Florian Metze. Mul- timodal abstractive summarization for How2 videos. In Anna Korhonen, David Traum, and Lluís Màrquez, editors, Proceedings of the Annual Meeting of the As- sociation for Computational Linguistics , pag...
2019
-
[136]
Lib- rispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Lib- rispeech: An ASR corpus based on public domain audio books. In Proceedings of 99 the IEEE International Conference on Acoustics, Speech and Signal Processing , pages 5206–5210, 2015
2015
-
[137]
Mindstorms: Children, computers, and powerful ideas
Seymour Papert. Mindstorms: Children, computers, and powerful ideas . Basic Books, 1980
1980
-
[138]
Pastore and Albert D
Raymond S. Pastore and Albert D. Ritzhaupt. Using time-compression to make multimedia learning more efficient: Current research and practice. TechTrends, 59:66–74, 2015
2015
-
[139]
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Y ang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...
2019
-
[140]
Video di- gests: A browsable, skimmable format for informational lecture videos
Amy Pavel, Colorado Reed, Björn Hartmann, and Maneesh Agrawala. Video di- gests: A browsable, skimmable format for informational lecture videos. In Pro- ceedings of the 27th Annual ACM Symposium on User Interface Software and Technology, pages 573–582. Association for Computin...
2014
-
[141]
Tuned models of peer assessment in MOOCs
Chris Piech, Jonathan Huang, Zhenghao Chen, Chuong Do, Andrew Ng, and Daphne Koller. Tuned models of peer assessment in MOOCs. In Proceedings of the Educational Data Mining , pages 153–160, 2013
2013
-
[142]
LLMs in education: Evaluation GPT and BERT models in student comment classification
Anabel Pilicita and Enrique Barra. LLMs in education: Evaluation GPT and BERT models in student comment classification. Multimodal Technologies and Interaction, 9(5):44, 2025
2025
-
[143]
Plass, Bruce D
Jan L. Plass, Bruce D. Homer, and Charles K. Kinzer. Foundations of game-based learning. Educational Psychologist, 50(4):258–283, 2015
2015
-
[144]
The tacit dimension
Michael Polanyi. The tacit dimension . Doubleday, 1966
1966
-
[145]
Content- aware dynamic timeline for video browse
Suporn Pongnumkul, Jue Wang, Gonzalo Ramos, and Michael Cohen. Content- aware dynamic timeline for video browse. In Proceedings of the annual ACM symposium on User interface software and technology , pages 139–142, 2010
2010
-
[146]
Discriminatively trained acoustic models for improving mispronunciation detection and diagnosis in computer aided pronunciation training (CAPT)
Xiaojun Qian, Frank Soong, and Helen Meng. Discriminatively trained acoustic models for improving mispronunciation detection and diagnosis in computer aided pronunciation training (CAPT). In Proceedings of INTERSPEECH, 2010
2010
-
[147]
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, p...
2023
-
[148]
Improv- ing language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improv- ing language understanding by generative pre-training. Technical report, OpenAI, 2018
2018
-
[149]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. Technical Re- port TR-2019-1, OpenAI, 2019. 100
2019
-
[150]
Radwan, Nancy M
Nisreen I. Radwan, Nancy M. Salem, and Mohamed I. El Adawy. Histogram correlation for video scene change detection. In Proceedings of the Second Inter- national Conference on Computer Science, Engineering and Applications , pages 765–773, 2012
2012
-
[151]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Y anqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020
2020
-
[152]
Y amamoto Ravenor
R. Y amamoto Ravenor. On the adequacy of Elsa Speak in formal education: A survey of teacher-users. Advances in Artificial Intelligence and Machine Learning, 04:2387–2394, 2024
2024
-
[153]
A theory of justice
John Rawls. A theory of justice . Harvard University Press, 1971
1971
-
[154]
ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope
Partha Pratim Ray. ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3:121–154, 2023
2023
-
[155]
Brian J. Reiser. Scaffolding complex learning: The mechanisms of structuring and problematizing student work. The Journal of the Learning Sciences , 13(3):273– 304, 2004
2004
-
[156]
Faster R-CNN: To- wards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: To- wards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors,Proceedings of the Ad- vances in Neural Information Processing Syste...
2015
-
[157]
Scratch: Programming for all
Mitchel Resnick, John Maloney, Andrés Monroy-Hernández, Natalie Rusk, Eve- lyn Eastmond, Karen Brennan, Amon Millner, Eric Rosenbaum, Jay Silver, Brian Silverman, and Y asmin Kafai. Scratch: Programming for all. Communications of the ACM, 52(11):60–67, 2009
2009
-
[158]
Risko, Nicola Anderson, Amrita Sarwal, Megan Engelhardt, and Alan Kingstone
Evan F. Risko, Nicola Anderson, Amrita Sarwal, Megan Engelhardt, and Alan Kingstone. Everyday attention: Variation in mind wandering and memory in a lecture. Applied Cognitive Psychology, 26(2):234–242, 2012
2012
-
[159]
Rogerson-Revell
Pamela M. Rogerson-Revell. Computer-assisted pronunciation training (CAPT): Current issues and future directions. RELC Journal, 52(1):189–205, 2021
2021
-
[160]
Tradition or innovation: A comparison of modern ASR methods for forced alignment
Rotem Rousso, Eyal Cohen, Joseph Keshet, and Eleanor Chodroff. Tradition or innovation: A comparison of modern ASR methods for forced alignment. In Pro- ceedings of INTERSPEECH 2024 , pages 1525–1529, 2024
2024
-
[161]
Chat- GPT in lesson preparation: A teacher choices trial
Palak Roy, Helen Poet, Ruth Staunton, Katherine Aston, and David Thomas. Chat- GPT in lesson preparation: A teacher choices trial. Technical report, National Foundation for Educational Research, December 2024
2024
-
[162]
Effects of second language pronunciation teach- ing revisited: A proposed measurement framework and meta-analysis
Kazuya Saito and Luke Plonsky. Effects of second language pronunciation teach- ing revisited: A proposed measurement framework and meta-analysis. Language Learning, 69(3):652–708, 2019
2019
-
[163]
Scheirer and M
E. Scheirer and M. Slaney. Construction and evaluation of a robust multifeature speech/music discriminator. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , volume 2, pages 1331–1334, 1997. 101
1997
-
[164]
wav2vec: Unsupervised pre-training for speech recognition.Proceedings of INTERSPEECH, pages 3465–3469, 2019
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised pre-training for speech recognition.Proceedings of INTERSPEECH, pages 3465–3469, 2019
2019
-
[165]
Generative AI in education: Past, present, and future
Tony Sheehan. Generative AI in education: Past, present, and future. EDUCAUSE Review, 2023
2023
-
[166]
Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Y ang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Y ang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R. J. Skerrv-Ryan, Rif A. Saurous, Y annis Agiomyrgiannakis, and Y onghui Wu. Natural TTS syn- thesis by conditioning Wavenet on MEL spectrogram predi...
2018
-
[167]
Sherwood
Bruce A. Sherwood. The computer speaks. IEEE Spectrum, 16(8):18–25, 1979
1979
-
[168]
Direct manipulation: A step beyond programming languages
Ben Shneiderman. Direct manipulation: A step beyond programming languages. Computer, 16(8):57–69, 1983
1983
-
[169]
Direct manipulation for comprehensible, predictable and con- trollable user interfaces
Ben Shneiderman. Direct manipulation for comprehensible, predictable and con- trollable user interfaces. In Proceedings of the International Conference on Intel- ligent User Interfaces, pages 33–39, 1997
1997
-
[170]
Human-centered AI
Ben Shneiderman. Human-centered AI. Oxford University Press, 2022
2022
-
[171]
Direct manipulation versus interface agents
Ben Shneiderman and Pattie Maes. Direct manipulation versus interface agents. Interactions, 4(6):42–61, 1997
1997
-
[172]
Gamification in mobile-assisted language learning: A sys- tematic review of Duolingo literature from public release of 2012 to early 2020
Mitchell Shortt, Shantanu Tilak, Irina Kuznetcova, Bethany Martens, and Ba- batunde Akinkuolie. Gamification in mobile-assisted language learning: A sys- tematic review of Duolingo literature from public release of 2012 to early 2020. Computer Assisted Language Learning, 36(3)...
2012
-
[173]
Augmented visual, auditory, haptic and multimodal feedback in motor learning: A review
Roland Sigrist, Georg Rauter, Robert Riener, and Peter Wolf. Augmented visual, auditory, haptic and multimodal feedback in motor learning: A review. Psycho- nomic Bulletin & Review , 20:21–53, 2013
2013
-
[174]
Herbert A. Simon. Designing organizations for an information-rich world. In Martin Greenberger, editor, Computers, Communications, and the Public Interest, pages 37–72. Johns Hopkins Press, 1971
1971
-
[175]
The critical period hypothesis for L2 acquisition: An unfalsifiable embarrassment? Languages, 6(3):149, 2021
David Singleton and Justyna Leśniewska. The critical period hypothesis for L2 acquisition: An unfalsifiable embarrassment? Languages, 6(3):149, 2021
2021
-
[176]
R. Smith. An overview of the Tesseract OCR Engine. In Proceedings of the International Conference on Document Analysis and Recognition, volume 2, pages 629–633, 2007
2007
-
[177]
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research , 15(56):1929–1958, 2014
1929
-
[178]
Sullivan, Shailesh S
Katherine J. Sullivan, Shailesh S. Kantak, and Patricia A. Burtner. Motor learning in children: Feedback effects on skill acquisition. Physical Therapy, 88(6):720– 732, 2008. 102
2008
-
[179]
Sutherland
Ivan E. Sutherland. Sketchpad: a man-machine graphical communication system. In Proceedings of the May 21-23, 1963, Spring Joint Computer Conference, pages 329–346. ACM, 1963
1963
-
[180]
Cognitive load during problem solving: Effects on learning
John Sweller. Cognitive load during problem solving: Effects on learning. Cog- nitive Science, 12(2):257–285, 1988
1988
-
[181]
Cognitive load theory
John Sweller. Cognitive load theory. In Psychology of learning and motivation , volume 55, pages 37–76. Elsevier, 2011
2011
-
[182]
John Sweller, Jeroen J. G. van Merriënboer, and Fred Paas. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31:261– 292, 2019
2019
-
[183]
The politics of recognition
Charles Taylor. The politics of recognition. In Amy Gutmann, editor, Multicul- turalism, pages 25–74. Princeton University Press, 1994
1994
-
[184]
Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers
Cristian Tejedor-García, Valentín Cardeñoso Payo, and David Escudero-Mancebo. Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers. Applied Sciences, 11(15), 2021
2021
-
[185]
Theepan Tharumalingam, Brady R. T. Roberts, Jonathan M. Fawcett, and Evan F. Risko. Increasing video lecture playback speed can impair test performance—a meta-analysis. Educational Psychology Review, 37, 2025
2025
-
[186]
How Khan Academy is changing the rules of education
Clive Thompson. How Khan Academy is changing the rules of education. Wired, July 2011
2011
-
[187]
Ron I. Thomson. Measurement of accentedness, intelligibility and comprehensi- bility. In Okim Kang and April Ginther, editors, Assessment in second language pronunciation, pages 11–29. Routledge, 2018
2018
-
[188]
EduQate: Generating adap- tive curricula through RMABs in education settings
Sidney Tio, Dexun Li, and Pradeep Varakantham. EduQate: Generating adap- tive curricula through RMABs in education settings. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems , pages 2042–2050, 2025. ACM Digital Library identifier ...
2025
-
[189]
Video abstraction: A systematic review and classification
Ba Tu Truong and Svetha Venkatesh. Video abstraction: A systematic review and classification. ACM Transactions on Multimedia Computing, Communications, and Applications, 3(1), 2007
2007
-
[190]
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. InProceedings of the ISCA Workshop on Speech Synthesis Workshop, page 125, 2016
2016
-
[191]
Audio summarization for podcasts
Aneesh Vartakavi, Amanmeet Garg, and Zafar Rafii. Audio summarization for podcasts. In Proceedings of the European Signal Processing Conference , pages 431–435. IEEE, 2021
2021
-
[192]
Gomez, Ł ukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems , vol- ume 30, pages 5998–6008, 2017. 103
2017
-
[193]
Vygotsky
Lev S. Vygotsky. Mind in society: The development of higher psychological pro- cesses. Harvard University Press, 1978
1978
-
[194]
Web Audio API
W3C Audio Working Group. Web Audio API. W3C Recommendation, 2021. 17 June 2021
2021
-
[195]
WebRTC 1.0: Real-time communication between browsers
W3C Web Real-Time Communications Working Group. WebRTC 1.0: Real-time communication between browsers. W3C Recommendation, 2021. 26 January 2021
2021
-
[196]
Fairseq S2T: Fast speech-to-text modeling with Fairseq
Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. Fairseq S2T: Fast speech-to-text modeling with Fairseq. In Proceedings of the Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Con...
2020
-
[197]
Artificial-intelligence-generated content with diffusion models: A literature review
Xiaolong Wang, Zhijian He, and Xiaojiang Peng. Artificial-intelligence-generated content with diffusion models: A literature review. Mathematics, 12(7):977, 2024
2024
-
[198]
Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Y onghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Y ang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc V . Le, Y annis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous. Tacotron: Towards end-to-end speech synthesis. In Proceedings o...
2017
-
[199]
Six views of embodied cognition
Margaret Wilson. Six views of embodied cognition. Psychonomic Bulletin & Review, 9:625–636, 2002
2002
-
[200]
Witt and Steve J
Silke M. Witt and Steve J. Y oung. Phone-level pronunciation scoring and assess- ment for interactive language learning. Speech Communication , 30(2):95–108, 2000
2000
-
[201]
Bruner, and Gail Ross
David Wood, Jerome S. Bruner, and Gail Ross. The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry , 17(2):89–100, 1976
1976
-
[202]
David R. Woolley. PLATO: The emergence of on-line community. Computer- Mediated Communication Magazine, 1(3):5, 1994
1994
-
[203]
Improving student-AI interaction through ped- agogical prompting: An example in computer science education
Ruiwei Xiao, Xinying Hou, Runlong Y e, Majeed Kazemitabaar, Nicholas Diana, Michael Liut, and John Stamper. Improving student-AI interaction through ped- agogical prompting: An example in computer science education. arXiv preprint arXiv:2506.19107v1, 2025. Submitted 23 June 2025
2025 arXiv
-
[204]
Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig
Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Michael L. Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig. Toward human parity in con- versational speech recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, pages 2410–2423, 2017
2017
-
[205]
Conformer-based speech recognition on extreme edge-computing devices
Mingbin Xu, Alex Jin, Sicheng Wang, Mu Su, Tim Ng, Henry Mason, Shiyi Han, Zhihong Lei, Y aqiao Deng, Zhen Huang, and Mahesh Krishnamoorthy. Conformer-based speech recognition on extreme edge-computing devices. In Pro- ceedings of the Conference of the North American Chapter o...
2024
-
[206]
VidChapters-7M: Video chapters at scale
Antoine Y ang, Arsha Nagrani, Ivan Laptev, Josef Sivic, and Cordelia Schmid. VidChapters-7M: Video chapters at scale. arXiv, 2023. arXiv:2309.13952. NeurIPS 2023 Datasets and Benchmarks Track
2023 arXiv
-
[207]
Re-examining whether, why, and how human-AI interaction is uniquely difficult to design
Qian Y ang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. Re-examining whether, why, and how human-AI interaction is uniquely difficult to design. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1–13, 2020
2020
-
[208]
SoftVideo: Improving the learning experience of software tutorial videos with collective interaction data
Saelyne Y ang, Jisu Yim, Aitolkyn Baigutanova, Seoyoung Kim, Minsuk Chang, and Juho Kim. SoftVideo: Improving the learning experience of software tutorial videos with collective interaction data. In Proceedings of the International Con- ference on Intelligent User Interfaces, ...
2022
-
[209]
Lin, Andy T
Shu-wen Y ang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakho- tia, Yist Y . Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu- Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahm...
2021
-
[210]
Personalized video summarization based on behavior of viewer
Atsuo Y oshitaka and Kazuya Sawada. Personalized video summarization based on behavior of viewer. In Proceedings of the International Conference on Signal Image Technology and Internet Based Systems, pages 661–667, 2012
2012
-
[211]
Predictive video analytics in online courses: A systematic literature review.Technology, Knowledge and Learning, 29:1907–1937, 2024
Ozan Raşit Yürüm, Tuğba Taşkaya-Temizel, and Soner Yıldırım. Predictive video analytics in online courses: A systematic literature review.Technology, Knowledge and Learning, 29:1907–1937, 2024
1907
-
[212]
Zekveld, Sophia E
Adriana A. Zekveld, Sophia E. Kramer, and Joost M. Festen. Cognitive load during speech perception in noise: The influence of age, hearing loss, and cognition on the pupil response. Ear and Hearing , 32(4):498–510, 2011
2011
-
[213]
Weiss, Y e Jia, Zhifeng Chen, and Y onghui Wu
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Y e Jia, Zhifeng Chen, and Y onghui Wu. LibriTTS: A corpus derived from LibriSpeech for text-to-speech. In Proceedings of INTERSPEECH, pages 1526–1530, 2019
2019
-
[214]
Instruc- tional video in e-learning: Assessing the impact of interactive video on learning effectiveness
Dongsong Zhang, Lina Zhou, Robert O Briggs, and Jay F Nunamaker Jr. Instruc- tional video in e-learning: Assessing the impact of interactive video on learning effectiveness. Information & Management, 43(1):15–27, 2006
2006
-
[215]
Video features, engagement, and patterns of collective attention allocation: An open flow network perspective
Jingjing Zhang, Yicheng Huang, and Ming Gao. Video features, engagement, and patterns of collective attention allocation: An open flow network perspective. Journal of Learning Analytics , 9(1):32–52, 2022
2022
-
[216]
WithY ou: Automated adaptive speech tutoring with context-dependent speech recognition
Xinlei Zhang, Takashi Miyaki, and Jun Rekimoto. WithY ou: Automated adaptive speech tutoring with context-dependent speech recognition. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1–2, 2020
2020
-
[217]
Embodied music training can help improve speech imitation and pronunciation skills.Language Teaching, 2024
Yuan Zhang, Florence Baills, and Pilar Prieto. Embodied music training can help improve speech imitation and pronunciation skills.Language Teaching, 2024. Pub- lished online 18 December 2024
2024
-
[218]
L2-Arctic: A non-native 105 English speech corpus
Guanlong Zhao, Evgeny Chukharev-Hudilainen, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Ricardo Gutierrez-Osuna, and John Levis. L2-Arctic: A non-native 105 English speech corpus. In Proceedings of Interspeech 2018 , pages 2783–2787, 2018
2018
-
[219]
The crowd in MOOCs: A study of learning patterns at scale
Xin Zhou, Aixin Sun, Jie Zhang, and Donghui Lin. The crowd in MOOCs: A study of learning patterns at scale. Interactive Learning Environments , 33(3):2136– 2150, 2025
2025
-
[220]
Zimmerman
Barry J. Zimmerman. Becoming a self-regulated learner: An overview. Theory into Practice, 41(2):64–70, 2002
2002
-
[221]
Effects of playback speed and language proficiency on listening comprehension of multilingual English learners
Jiaxuan Zong, Nihat Polat, and Laura Mahalingappa. Effects of playback speed and language proficiency on listening comprehension of multilingual English learners. International Multilingual Research Journal , 2024. Published online 8 June 2024; later assigned to volume 19, iss...
2024
-
[222]
DDSupport: Language Learning Support System that Displays Differences and Distances from Model Speech,
Kazuki Kawamura and Jun Rekimoto, “DDSupport: Language Learning Support System that Displays Differences and Distances from Model Speech,” 21st IEEE International Conference on Machine Learning and Applications (ICMLA), Nas- sau, Bahamas, 2022, pp. 313-320, doi: 10.1109/ICMLA5...
2022
-
[223]
AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,
Kazuki Kawamura and Jun Rekimoto, “AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,” Augmented Hu- mans International Conference 2023 (AHs), Glasgow, United Kingdom, 2023, pp. 200-208, doi: 10.1145/3582700.3582722
2023
-
[224]
FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Con- texts,
Kazuki Kawamura and Jun Rekimoto, “FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Con- texts,” Augmented Humans International Conference 2024 (AHs), Melbourne, Aus- tralia, 2024, pp. 205-216, doi: 10.1145/3652920...
2024
-
[225]
A Language Acquisition Support System that Presents Differences and Distances from Model Speech,
Kazuki Kawamura and Jun Rekimoto, “A Language Acquisition Support System that Presents Differences and Distances from Model Speech,” 34th Annual ACM Symposium on User Interface Software and Technology (UIST), Virtual Event, USA, 2021, pp. 44-46, doi: 10.1145/3474349.3480225
2021
-
[226]
Visualization of Speech Differences for Dialect Speech Training,
Kazuki Kawamura and Jun Rekimoto, “Visualization of Speech Differences for Dialect Speech Training,” 2021 CHI Conference on Human Factors in Comput- ing Systems Workshop on Human Augmentation for Skill Acquisition and Skill Transfer (CHI), Virtual Event, Japan, 2021
2021
-
[227]
AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,
Kazuki Kawamura and Jun Rekimoto, “AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,” 35th Annual ACM Symposium on User Interface Software and Technology (UIST), Bend, USA, 2022, pp. 1-3, doi: 10.1145/3526114.3558727
2022
-
[228]
QA-FastPerson: Extending Video Platform Search Capabilities by Creating Summary Videos in Response to User Queries,
Kazuki Kawamura and Jun Rekimoto, “QA-FastPerson: Extending Video Platform Search Capabilities by Creating Summary Videos in Response to User Queries,” Augmented Humans International Conference 2024 (AHs), Melbourne, Australia, 2024, pp. 290-293, doi: 10.1145/3652920.3653052
2024
-
[229]
Generating Summary Videos from User Questions to Support Video-Based Learning,
Kazuki Kawamura and Jun Rekimoto, “Generating Summary Videos from User Questions to Support Video-Based Learning,” 2024 CHI Conference on Human Factors in Computing Systems Workshop on Generative AI and HCI (CHI), Hon- olulu, USA, 2024. 107 Peer-Reviewed Domestic Conference Pr...
2024
-
[230]
DDSupport: A Language Learning Sup- port System that Presents Differences and Distances from Model Pronunciation,
Kazuki Kawamura and Jun Rekimoto, “DDSupport: A Language Learning Sup- port System that Presents Differences and Distances from Model Pronunciation,” Interaction, 2022, pp. 77-86 (in Japanese)
2022
-
[231]
FastPerson: Lecture Video Summarization Based on Visual and Audio Information for User-Centered Learning Experience,
Kazuki Kawamura and Jun Rekimoto, “FastPerson: Lecture Video Summarization Based on Visual and Audio Information for User-Centered Learning Experience,” Interaction, 2024, pp. 11-20 (in Japanese) 108
2024
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.