Pith. sign in

REVIEW 4 major objections 6 minor 231 references

AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques

T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This dissertation claims that deep-learning systems at three learning stages—phoneme-level adaptive playback, multimodal video summarization, and unlabeled-data pronunciation feedback—can make audio-visual learning faster and more…

desk verdict A three-system dissertation with a genuinely novel AIxSpeed core, but the load-bearing evaluation of that core is partly circular and under-tested; it deserves review, not rejection. read the letter →

arxiv 2608.08990 v1 pith:56IP7E6O submitted 2026-08-10 cs.HC cs.MM

classification cs.HCcs.MM
keywords AI-guidedlearningadaptiveplaybackspeedASRconfidencephoneme-levelcontrolvideosummarizationvoicecloningpronunciationfeedbackself-supervisedspeech
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation takes on two problems that make audio and video frustrating learning media: consuming long-form content sequentially takes too long, and imitation-based practice lacks scalable feedback. It proposes a Consume–Understand–Imitate framework and evaluates one deep-learning system for each stage. The central empirical claims are that AIxSpeed achieves an average playback factor of about 1.3x at higher listener opinion scores than matched constant-speed audio, that FastPerson cuts lecture-video viewing time by 53% with no statistically significant quiz-score difference, and that Profy produces a larger observed pronunciation-intelligibility gain than imitation-only practice, with non-overlapping pre- and post-practice confidence intervals. The dissertation presents these results as establishing an integrated technical basis for AI-guided learning from audio-visual content. If the results hold, learners could reclaim a large fraction of listening and viewing time and receive localized practice feedback without measured loss of comprehension.

What carries the argument

The object carrying the speed argument is the shared Wav2Vec2-based network in AIxSpeed, trained with Loss = Lossspeed + λLossctc: one head outputs a per-segment playback factor and the other performs CTC speech recognition, so a single model decides how fast each phoneme can go and verifies that the accelerated speech stays recognizable. The object carrying the summarization argument is FastPerson's video-to-video pipeline: chapter segmentation from scene and silence detection, visual metadata from OCR and object detection, transcription via Whisper, LLM-generated summaries, and VITS voice cloning to keep the narrator's voice continuous with the original. The object carrying the feedback argument is Profy's classifier, trained on unannotated learner and native speech, which scores utterances and visualizes the waveform regions the classifier attends to plus distances in latent space.

What would settle it

Take the same sentence materials from LibriSpeech and UME-ERJ, run AIxSpeed's per-phoneme speed profile, and have native and non-native listeners transcribe the output at word or phoneme level. If human accuracy does not track ASR confidence in those phoneme-level segments, or if the r≈0.998 correlation from the sentence-level pilot drops for non-native speech, then the claimed intelligibility preservation is unsupported. A second check: a quiz-based replication of FastPerson on unseen lectures with summary-only viewing would test whether the 53% time saving hides comprehension differences beyond the two videos used.

Watch

Extended reading notes

Core claim

The central claim is that speech-recognition confidence is a usable proxy for human listening difficulty at speeds above 1.0x, and that this proxy, a multimodal summarizer with voice preservation, and a self-supervised proficiency model each carry their respective stage of the learning cycle. The pilot study reports a correlation of r=0.9977 between human transcription accuracy and ASR accuracy across the tested speeds, and the full system is trained with a combined loss that maximizes speed while minimizing CTC recognition error, giving average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ with higher mean opinion scores than matched constant-speed playback. FastPerson is claimed to reduce viewing time by 53% on educational videos with quiz scores statistically indistinguishable from normal playback, and Profy is claimed to improve pronunciation intelligibility with pre- and post-practice confidence intervals that do not overlap the elicited-imitation baseline. The paper frames these three results as components of an integrated AI-Guided Learning loop rather than as standalone tools.

Load-bearing premise

The speed-optimization system stands or falls on the assumption that a speech recognizer's confidence remains a reliable proxy for human listening difficulty when applied at phoneme granularity to speech outside the tested English corpora, including non-native speech; the paper validates the proxy only at coarser granularity and for the specific datasets.

Editorial extensions

If this is right

  • Learners using AIxSpeed would get roughly 23% shorter listening sessions at the same average quality, since 1.3x playback turns a 60-minute lecture into about 46 minutes.
  • Video platforms could offer per-chapter summary versions that halve viewing time while quiz-measured comprehension stays statistically unchanged, with learners able to drill into full chapters when needed.
  • Pronunciation practice with model-derived localized feedback can improve intelligibility without requiring transcribed error annotations, lowering the cost of building feedback systems for new languages or skills.
  • The three systems compose a single learning loop: consume faster, understand via summaries or full segments, imitate with feedback, then return to relevant segments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct consequence the author leaves implicit is that AIxSpeed's speed profile could be personalized by raising or lowering the intelligibility threshold; the paper notes personalization as future work but does not test it.
  • Because the pilot correlation was measured at utterance and sentence level, the strongest untested extension is the assumption that the same proxy holds at phoneme boundaries; a per-phoneme human transcription study would settle it.
  • FastPerson's fixed summary-length weights (0.3 for audio, 2.5 for visual) suggest a testable extension: adapting those weights to learner proficiency or content type could yield further time savings or comprehension gains, which the dissertation does not test.
  • The same Profy mechanism—a classifier trained on unannotated expert versus learner data with region highlighting—could transfer to other imitation domains such as music or movement, though the paper evaluates only English pronunciation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The dissertation proposes an AI-guided learning framework organized around three stages (Consume, Understand, Imitate) and instantiates it in three systems: AIxSpeed, which adapts audio playback speed at the phoneme level using speech-recognition confidence as a proxy for listening difficulty; FastPerson, which creates chapter-based video summaries preserving both visual and spoken information and uses voice cloning for continuity; and Profy, which provides pronunciation feedback from largely unannotated audio by highlighting classifier-emphasized waveform regions and latent-space distances. Reported evaluations include: AIxSpeed achieving average playback factors of 1.30x (LibriSpeech) and 1.29x (UME-ERJ) with higher mean opinion scores than matched constant-speed playback; FastPerson reducing viewing time by 53% with no statistically significant quiz-score difference; and Profy showing observed intelligibility improvements with non-overlapping pre/post confidence intervals.

Significance. If the central claims held, the work would make a useful contribution to human-computer interaction and educational technology: it demonstrates concrete, deployable prototypes that address real bottlenecks in audio-visual learning, and it integrates self-supervised speech representations, LLM-based multimodal summarization, and low-resource pronunciation feedback. The manuscript is unusually transparent: it repeatedly states where evidence is missing, e.g., that the AIxSpeed MOS differences were not significance-tested, that the UME-ERJ result does not establish human intelligibility, and that Profy's highlighted regions were not independently validated as phonetic error diagnoses. That transparency is a strength, but it also means the abstract and chapter-level summaries frequently assert more than the evidence supports. The framework itself (Consume-Understand-Imitate) is a reasonable organizing device, and each prototype addresses a genuine gap, so the work has clear value as a systems/HCI contribution if the claims are appropriately qualified.

major comments (4)
  1. [§3.5.1 and §3.3.2] The technical evaluation of AIxSpeed is partly self-referential. The playback speed adjuster is trained jointly with the speech recognizer through the composite loss Loss = Lossspeed + λLossctc (§3.3.2), and the CER/WER evaluation in §3.5.1 uses "the speech recognizer used in our method" to score the modified audio. The lower CER/WER relative to constant-speed playback therefore reflects, at least in part, that the speed profile was optimized to be recognizable to this exact evaluator. The pilot survey in §3.2 validates only aggregate, whole-utterance correlation at constant speeds; it does not validate per-phoneme confidence as a proxy for human intelligibility on non-native or time-stretched speech. The claim that intelligibility is preserved requires either an independent recognizer of a different architecture or a direct human transcription test at phoneme/segment granularity.
  2. [§3.5.2] The higher mean opinion scores for AIxSpeed over matched constant-speed playback are reported without a significance test; the text itself states that "The source paper did not report a significance test for these differences." A 0.5-point and 0.8-point difference on a five-point MOS scale is not interpretable without a test or confidence interval, especially with 50 participants and 40 sentences. Additionally, MOS is a measure of perceived quality, not comprehension, so this result cannot by itself support the abstract's implication that listening comprehension is maintained.
  3. [§4.4.3] The FastPerson quiz result is internally consistent in reporting no statistically significant difference, but the conclusion that comprehension was retained is stronger than the evidence warrants. The comparison is between 19 and 21 participants, and the quiz-score standard deviations are large (e.g., Video 2: 0.67 ± 0.46 for FastPerson vs. 0.63 ± 0.39 for control). A non-significant t-test with these sample sizes does not establish equivalence; a non-inferiority test or a confidence interval for the quiz-score difference is needed. The viewing-time reduction (53%) is well supported by the reported t-tests, so the efficiency claim is robust; the claim of preserved understanding is not established at the same evidentiary level.
  4. [§5.4 and §5.5.2] The Profy evaluation involves only 10 learners and 5 raters, and the headline finding is presented as non-overlapping pre- and post-practice confidence intervals without reporting the interval values or a test statistic. The observed intelligibility gain could reflect practice effects or rater drift, and the elicited-imitation baseline is not accompanied by effect sizes. Furthermore, as the manuscript itself acknowledges in §2.8.2 and §5.5.2, the assumption that classifier-emphasized waveform regions and latent distances correspond to correctable pronunciation deviations was not independently validated. The chapter's central claim of improved pronunciation intelligibility is therefore plausible but not yet firmly supported; reporting raw scores, CIs, and an effect size would help, as would a control condition for the feedback mechanism itself.
minor comments (6)
  1. [§3.2] The formulas for WER and CER state the denominator as "number of correct words" (and characters); standard definitions divide by the total number of reference words/characters. As written, the metric is non-standard and should be clarified or corrected.
  2. [Table 3.1] The CER/WER comparisons in Table 3.1 report point estimates only; confidence intervals or significance tests for the differences would improve interpretability.
  3. [Abstract and §1.7.2] The abstract and the contribution list state that AIxSpeed "received higher mean opinion scores" and that UME-ERJ "suggests" improved listenability for non-native speech. Given the explicit caveats in §3.5.2 and §3.6.4, these statements should be qualified as observed differences without significance testing.
  4. [§4.4.3 and Figure 4.3] The text refers to the "right side" and "left side" of Figure 4.3, but the figure appears as a single panel of per-question correct-answer rates. Please align the description with the actual figure layout.
  5. [§4.2.4] The summary-length weights ws = 0.3 and wv = 2.5 are described as determined by preliminary experiments; reporting the criterion used (e.g., which pilot data, what target) would improve reproducibility.
  6. [§5.4] For the Profy result, please report the actual confidence interval values rather than only stating that they do not overlap, and specify the statistical procedure used to construct them.

Circularity Check

1 steps flagged · score 6.0 of 10

AIxSpeed's ASR-intelligibility result is circular: the speed adjuster is trained against the same recognizer that is then used to evaluate it, leaving only an un-significance-tested MOS comparison as external support.

  1. fitted input called prediction [Section 3.3.2 (Speech Recognizer), Section 3.4.1 (Implementation Details), Section 3.5.1 (Technical Evaluation)]
    "In summary, the entire model is trained to minimize the following error function, which is a combination of this error function and the error function of the playback speed adjuster: Loss = Lossspeed + λLossctc. ... A standard Wav2Vec2-based speech recognition model, which was the speech recognizer used in our method, was used to compute the CER and WER for comparison. ... This indicates that it was more recognizable to the ASR model used for evaluation; it does not directly establish improved intelligibility or comprehension for human listeners, which requires a dedicated human study."

    The playback-speed adjuster is trained with a joint objective that includes the recognizer's CTC loss (Loss = Lossspeed + λLossctc), and the recognizer's weights are held fixed while the adjuster is optimized against that recognizer. The technical evaluation then measures CER/WER with 'the speech recognizer used in our method'—the exact model whose loss was part of the training objective. Lower CER/WER for AIxSpeed than for matched constant-speed playback therefore partly reflects that the adjuster was fitted to make this recognizer succeed, not an independent measure of human intelligibility.

full rationale

The clearest circularity is localized to AIxSpeed. The speed adjuster is optimized with a loss that explicitly includes the speech recognizer's CTC loss, and the technical evaluation uses that same recognizer to compute CER/WER; the observed improvement over matched constant-speed playback is thus partly a self-consistency check of the training objective rather than an independent intelligibility result. The paper is transparent about this limitation in the quoted passages, but the chapter conclusion still leans on the ASR numbers to claim intelligibility preservation. The pilot survey correlating human and ASR transcription accuracy (r=0.9977) is independent external evidence, though it validates only aggregate speed conditions and not per-phoneme decisions. FastPerson's viewing-time reduction is a measured outcome rather than a fitted prediction, and its quiz scores are independent human outcomes; Profy is evaluated by external human raters. There is no load-bearing self-citation chain. Overall, one central technical result reduces by construction to its training objective, warranting a partial-circularity score of 6.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

The central claims rest on several domain assumptions: ASR confidence standing in for human listening difficulty, LLM summaries preserving pedagogical content, voice cloning preserving learning continuity, and classifier-localized regions corresponding to pronunciation errors. The free parameters listed are hand-set or fitted in preliminary experiments. No new physical entities are introduced.

free parameters (4)
  • AIxSpeed loss weight lambda = 1e-7
    Balances speed maximization and CTC recognition in the combined loss; fixed by hand and not tuned on an independent validation set.
  • FastPerson summary length weights ws and wv = ws=0.3, wv=2.5
    Used in N = max(Nmin, ws*Lt + wv*(Lo+Lc)) to set summary length; calibrated to pilot lecture videos and not cross-validated.
  • FastPerson segmentation thresholds = Bhattacharyya 0.6; silence RMS below 5% for >1s
    Scene and silence detection thresholds empirically determined in preliminary experiments.
  • FastPerson minimum summary length Nmin = 50 words
    Chosen to avoid over-terse summaries; no independent justification is given.
assumptions (6)
  • domain assumption ASR confidence is a valid proxy for human listening difficulty for the purpose of setting playback speed.
    Pilot shows aggregate correlation 0.9977 for speeds above 1.0x, but the system applies it at phoneme level and to non-native speech without per-phoneme human validation (Sections 3.2, 3.3).
  • domain assumption The aggregate human-ASR correlation extrapolates to fine-grained per-phoneme speed decisions.
    The pilot compared whole-sentence transcription at constant speeds; AIxSpeed varies speed by phoneme, a regime not directly validated against human listening (Section 3.3).
  • domain assumption LLM-based summaries from transcripts, OCR, and object labels retain the content needed for quiz performance.
    FastPerson relies on GPT-3.5 to distill multimodal inputs; summary quality is not independently evaluated except through the final quiz (Section 4.2.3).
  • domain assumption Voice cloning preserves sufficient speaker identity and prosody for seamless learning continuity.
    VITS speaker adaptation is used to generate narration; no direct perception test of voice similarity is reported (Section 4.2.5).
  • ad hoc to paper Classifier-emphasized waveform regions and latent distances correspond to correctable pronunciation deviations.
    Profy's feedback assumes model attention localizes errors; the dissertation explicitly notes this was not independently validated as phonetic error diagnosis (Section 2.8.2).
  • domain assumption Absence of a statistically significant quiz-score difference in a small sample is treated as evidence that summarized viewing does not harm learning.
    FastPerson compared 19 vs 21 participants; the study may be underpowered for a non-inferiority conclusion (Section 4.4.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques." pith.science (2026). https://pith.science/paper/56IP7E6O

@misc{pith2026260808990,
  author       = {Pith},
  title        = {Pith review of: AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/56IP7E6O}},
  note         = {Machine review of arXiv:2608.08990}
}
read the original abstract

Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.

Figures

Figures reproduced from arXiv: 2608.08990 by the authors.

Figure 1.1
Figure 1.1. An example of a Massive Open Online Course (MOOC) interface, typically [PITH_FULL_IMAGE:figures/full_fig_p012_1_1.png] view at source ↗
Figure 1.2
Figure 1.2. Information-density variations inferred from collective click-frequency traces [PITH_FULL_IMAGE:figures/full_fig_p013_1_2.png] view at source ↗
Figure 1.3
Figure 1.3. An example of imitation learning in an online dance class. Learners attempt [PITH_FULL_IMAGE:figures/full_fig_p013_1_3.png] view at source ↗
Figures from the paper (23 more)
Figure 1.4
Figure 1.4. Figure 1.4: The iterative Consume‒Understand‒Imitate learning loop. Each stage is later augmented by a dedicated AI subsystem (AIxSpeed, FastPerson, and Profy; see §1.5.4). Furthermore, in asynchronous online learning environments, feedback delay is also an important issue. Lear…
Figure 2.1
Figure 2.1. Figure 2.1: A PLATO terminal in use (c. 1972). The networked PLATO system fea [PITH_FULL_IMAGE:figures/full_fig_p023_2_1.png]
Figure 2.2
Figure 2.2. Figure 2.2: The dynamic timeline approach decouples video speed from playback speed, [PITH_FULL_IMAGE:figures/full_fig_p025_2_2.png]
Figure 2.3
Figure 2.3. Figure 2.3: Dynamic skim generation using hierarchical content analysis. The system [PITH_FULL_IMAGE:figures/full_fig_p026_2_3.png]
Figure 3.1
Figure 3.1. Figure 3.1: AIxSpeed optimizes the playback speed of a video in units as small as [PITH_FULL_IMAGE:figures/full_fig_p041_3_1.png]
Figure 3.2
Figure 3.2. Figure 3.2: Comparison of playback speed and listening comprehension in machine [PITH_FULL_IMAGE:figures/full_fig_p043_3_2.png]
Figure 3.3
Figure 3.3. Figure 3.3: The work process of AIxSpeed: extracting human voices from the target [PITH_FULL_IMAGE:figures/full_fig_p045_3_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: AIxSpeed architecture: Our system simultaneously optimizes the playback [PITH_FULL_IMAGE:figures/full_fig_p046_3_4.png]
Figure 3.5
Figure 3.5. Figure 3.5: Application example showing how the playback speed changed when [PITH_FULL_IMAGE:figures/full_fig_p047_3_5.png]
Figure 3.6
Figure 3.6. Figure 3.6: Voice quality comparison between baseline (constant playback speed increase [PITH_FULL_IMAGE:figures/full_fig_p051_3_6.png]
Figure 4.1
Figure 4.1. Figure 4.1: FastPerson: Video summarization method that generates a summary sentence [PITH_FULL_IMAGE:figures/full_fig_p058_4_1.png]
Figure 4.2
Figure 4.2. Figure 4.2: Overview of the FastPerson user interface: A Video Window showcases [PITH_FULL_IMAGE:figures/full_fig_p062_4_2.png]
Figure 4.3
Figure 4.3. Figure 4.3: Correct answer rates for each question of the video. [PITH_FULL_IMAGE:figures/full_fig_p065_4_3.png]
Figure 4.4
Figure 4.4. Figure 4.4: Average video viewing time by viewing method. [PITH_FULL_IMAGE:figures/full_fig_p066_4_4.png]
Figure 4.5
Figure 4.5. Figure 4.5: Survey results on user experience. Q4: What improvements would you suggest for the user interface design? (Open￾ended Response) 2. Learning Experience Satisfaction Q5: Was learning with our system enjoyable? (1: Not Enjoyable At All - 5: Ex￾tremely Enjoyable) Q6: Com…
Figure 5.1
Figure 5.1. Figure 5.1: Profy architecture: Overview of the system’s components and data flow. [PITH_FULL_IMAGE:figures/full_fig_p076_5_1.png]
Figure 5.2
Figure 5.2. Figure 5.2: Profy pronunciation score determination: Process of classifying speech data [PITH_FULL_IMAGE:figures/full_fig_p077_5_2.png]
Figure 5.3
Figure 5.3. Figure 5.3: Profy difference visualization: Method for highlighting parts of speech that [PITH_FULL_IMAGE:figures/full_fig_p080_5_3.png]
Figure 5.4
Figure 5.4. Figure 5.4: Profy distance visualization: Representation of speech data in two [PITH_FULL_IMAGE:figures/full_fig_p081_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Profy user interface: Layout and components of the application’s main screen. [PITH_FULL_IMAGE:figures/full_fig_p082_5_5.png]
Figure 5.6
Figure 5.6. Figure 5.6: Examples of Profy interactions: Demonstration of eliminating differences [PITH_FULL_IMAGE:figures/full_fig_p082_5_6.png]
Figure 5.7
Figure 5.7. Figure 5.7: Comparison of post-practice intelligibility gains for users of Profy and stan [PITH_FULL_IMAGE:figures/full_fig_p084_5_7.png]
Figure 5.8
Figure 5.8. Figure 5.8: Questionnaire results: User feedback on Profy’s ease of use and helpfulness [PITH_FULL_IMAGE:figures/full_fig_p084_5_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

231 extracted references · 71 canonical work pages

  1. [1]

    The effect of settings, educational level and tools on computer-assisted pronun- ciation training: A meta-analysis

    Asma Almusharraf, Hassan Saleh Mahdi, Haifa Al-Nofaie, and Amal Aljasser. The effect of settings, educational level and tools on computer-assisted pronun- ciation training: A meta-analysis. Journal of Computer Assisted Learning , 40(4):1605–1615, 2024

  2. [2]

    Advances in computer-based education: The Plato program will provide a major test of the educational and economic fea- sibility of this medium

    Daniel Alpert and Donald Lester Bitzer. Advances in computer-based education: The Plato program will provide a major test of the educational and economic fea- sibility of this medium. Science, 167(3925):1582–1590, 1970

  3. [3]

    Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. Guidelines for human-AI interac- tion. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2019

  4. [4]

    Computer-assisted pronunciation training: A systematic review

    Moustafa Amrate and Pi-hua Tsai. Computer-assisted pronunciation training: A systematic review. ReCALL, 37(1):22–42, 2025. Published online 19 September 2024

  5. [5]

    Anderson and Brian J

    John R. Anderson and Brian J. Reiser. The LISP tutor. BYTE, 10(4):159–175, 1985

  6. [6]

    Anderson and David R

    Lorin W. Anderson and David R. Krathwohl, editors. A Taxonomy for Learning, Teaching, and Assessing: A Revision of Bloom’s Taxonomy of Educational Objec- tives. Longman, New Y ork, 2001

  7. [7]

    Metsai, Vasileios Mezaris, and Ioannis Patras

    Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, and Ioannis Patras. Video summarization using deep neural networks: A survey. Proceedings of the IEEE , 109(11):1838–1863, 2021

  8. [8]

    Common voice: A massively-multilingual speech corpus

    Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis T yers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 4218–4222. European Language Resources Association, 2020

Show all 231 references
  1. [9]

    Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Y ongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi

    Sercan Ö. Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Y ongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi. Deep voice: Real-time neural text- to-speech. In Proceedings of the International Conference...

  2. [10]

    ViViT: A video vision transformer

    Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. ViViT: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 6816–6826, 2021. 90

  3. [11]

    SpeechSkimmer: A system for interactively skimming recorded speech

    Barry Arons. SpeechSkimmer: A system for interactively skimming recorded speech. ACM Transactions on Computer-Human Interaction, 4(1):3–38, 1997

  4. [12]

    Cromley, and Diane Seibert

    Roger Azevedo, Jennifer G. Cromley, and Diane Seibert. Does adaptive scaffold- ing facilitate students’ ability to regulate their learning with hypermedia? Con- temporary Educational Psychology, 29(3):344–370, 2004

  5. [13]

    wav2vec 2.0: A framework for self-supervised learning of speech representations

    Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems, volume 33, pages 12449– 12460. Curran Associates, Inc., 2020

  6. [14]

    Neural machine transla- tion by jointly learning to align and translate

    Dzmitry Bahdanau, Kyunghyun Cho, and Y oshua Bengio. Neural machine transla- tion by jointly learning to align and translate. In Proceedings of the International Conference on Learning Representations, 2015

  7. [15]

    Speaker recognition based on deep learning: An overview

    Zhongxin Bai and Xiao-Lei Zhang. Speaker recognition based on deep learning: An overview. Neural Networks, 140:65–99, 2021

  8. [16]

    Social learning theory

    Albert Bandura. Social learning theory . Prentice Hall, 1977

  9. [17]

    Self-efficacy: The exercise of control

    Albert Bandura. Self-efficacy: The exercise of control . W. H. Freeman, 1997

  10. [18]

    Meeting the universe halfway: Quantum physics and the entangle- ment of matter and meaning

    Karen Barad. Meeting the universe halfway: Quantum physics and the entangle- ment of matter and meaning . Duke University Press, 2007

  11. [19]

    Text summarization using large language mod- els: A comparative study of MPT-7b-instruct, Falcon-7b-instruct, and OpenAI Chat-GPT models

    Lochan Basyal and Mihir Sanghvi. Text summarization using large language mod- els: A comparative study of MPT-7b-instruct, Falcon-7b-instruct, and OpenAI Chat-GPT models. arXiv, 2023

  12. [20]

    Enhancing learning ex- periences: EEG-based passive BCI system adapts learning speed to cognitive load in real-time, with motivation as catalyst

    Noémie Beauchemin, Patrick Charland, Alexander Karran, Jared Boasen, Bella Tadson, Sylvain Sénécal, and Pierre-Majorique Léger. Enhancing learning ex- periences: EEG-based passive BCI system adapts learning speed to cognitive load in real-time, with motivation as catalyst. Fro...

  13. [21]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the ACM Conference on Fairness, Accountability, and Trans- parency, pages 610–623. Association for Computing M...

  14. [22]

    Abstractive video lecture summarization: Applications and future prospects

    Irene Benedetto, Moreno La Quatra, Luca Cagliero, Lorenzo Canale, and Laura Farinetti. Abstractive video lecture summarization: Applications and future prospects. Education & Information Technologies, 29(3):2951–2971, 2024

  15. [23]

    The evidence for ’flipping out’: A systematic review of the flipped classroom in nursing education

    Vasiliki Betihavas, Heather Bridgman, Rachel Kornhaber, and Merylin Cross. The evidence for ’flipping out’: A systematic review of the flipped classroom in nursing education. Nurse Education Today, 38:15–21, 2016

  16. [24]

    Bitzer, Peter G

    Donald L. Bitzer, Peter G. Braunfeld, and Wayne W. Lichtenberger. PLATO: An automatic teaching device. IRE Transactions on Education, 4(4):157–161, 1961

  17. [25]

    Duolingo language report 2024, 2024

    Cindy Blanco. Duolingo language report 2024, 2024. Retrieved May 25, 2025, from https://blog.duolingo.com/2024-duolingo-language-report/

  18. [26]

    Benjamin S. Bloom. The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6):4– 16, 1984. 91

  19. [27]

    Outline of a theory of practice

    Pierre Bourdieu. Outline of a theory of practice . Cambridge University Press, 1977

  20. [28]

    Cynthia J. Brame. Effective educational videos: Principles and guidelines for maximizing student learning from video content. CBEŠLife Sciences Education, 15(4):es6, 2016

  21. [29]

    PTeacher: A computer-aided personalized pronunciation training system with exaggerated audio-visual corrective feedback

    Y aohua Bu, Tianyi Ma, Weijun Li, Hang Zhou, Jia Jia, Shengqi Chen, Kaiyuan Xu, Dachuan Shi, Haozhe Wu, Zhihan Y ang, Kun Li, Zhiyong Wu, Yuanchun Shi, Xi- aobo Lu, and Ziwei Liu. PTeacher: A computer-aided personalized pronunciation training system with exaggerated audio-visu...

  22. [30]

    XTTS: A massively multilingual zero-shot text-to-speech model

    Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Lo- gan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, and Ju- lian Weber. XTTS: A massively multilingual zero-shot text-to-speech model. In Proceedings of INTERSPEECH 2024 , pages 4978–...

  23. [31]

    Seamful and seamless design in ubiquitous computing

    Matthew Chalmers and Ian MacColl. Seamful and seamless design in ubiquitous computing. In Workshop on At the Crossroads: The Interaction of HCI and Sys- tems Issues in UbiComp (UbiComp 2003 Workshop) , 2003. Position paper

  24. [32]

    Kumar, Rhea Varkhedi, and Dillon H

    Ashley Chen, Suchita E. Kumar, Rhea Varkhedi, and Dillon H. Murphy. The effect of playback speed and distractions on the comprehension of audio and audio-visual materials. Educational Psychology Review, 36:79, 2024

  25. [33]

    Video browse - a study of user behavior in online VoD services

    Liang Chen, Yipeng Zhou, and Dah Ming Chiu. Video browse - a study of user behavior in online VoD services. In Proceedings of the International Conference on Computer Communication and Networks, pages 1–7, 2013

  26. [34]

    Computer assisted pronunciation training (CAPT): A systematic review of studies from 2012 to 2021

    Xu Chen, Jie Mu, and Tingting Zhang. Computer assisted pronunciation training (CAPT): A systematic review of studies from 2012 to 2021. In Proceedings of the International Conference on Computers in Education , pages 575–580. Asia- Pacific Society for Computers in Education, 2022

  27. [35]

    MultiPA: A multi-task speech pronunciation assessment model for open response scenarios

    Yu-Wen Chen, Zhou Yu, and Julia Hirschberg. MultiPA: A multi-task speech pronunciation assessment model for open response scenarios. In Proceedings of INTERSPEECH, pages 297–301, 2024

  28. [36]

    The ICAP framework: Linking cognitive engagement to active learning outcomes

    Michelene TH Chi and Ruth Wylie. The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4):219– 243, 2014

  29. [37]

    MixT: Automatic generation of step-by-step mixed media tutorials

    Pei- Yu Chi, Sally Ahn, Amanda Ren, Mira Dontcheva, Wilmot Li, and Björn Hart- mann. MixT: Automatic generation of step-by-step mixed media tutorials. In Proceedings of the Annual ACM Symposium on User Interface Software and Tech- nology, pages 93–102, 2012

  30. [38]

    VIVID: Human-AI col- laborative authoring of vicarious dialogues from lecture videos

    Seulgi Choi, Hyewon Lee, Y oonjoo Lee, and Juho Kim. VIVID: Human-AI col- laborative authoring of vicarious dialogues from lecture videos. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–26. As- sociation for Computing Machinery, 2024

  31. [39]

    Attention-based models for speech recognition

    Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, KyungHyun Cho, and Y oshua Bengio. Attention-based models for speech recognition. In Proceedings 92 of the International Conference on Neural Information Processing Systems, pages 577–585, 2015

  32. [40]

    The extended mind

    Andy Clark and David Chalmers. The extended mind. Analysis, 58(1):7–19, 1998

  33. [41]

    Reac- tive video: Adaptive video playback based on user motion for supporting physical activity

    Christopher Clarke, Larissa Pschetz, Mor Trope, and Dave Murray-Rust. Reac- tive video: Adaptive video playback based on user motion for supporting physical activity. In Proceedings of the International Conference on Intelligent User Inter- faces, pages 196–208, 2020

  34. [42]

    Measuring mind wandering during online lectures assessed with EEG

    Colin Conrad and Aaron Newman. Measuring mind wandering during online lectures assessed with EEG. Frontiers in Human Neuroscience, 15:697532, 2021

  35. [43]

    The effects of time-compressed speech on native and EFL listening comprehension

    Linda Conrad. The effects of time-compressed speech on native and EFL listening comprehension. Studies in Second Language Acquisition , 11(1):1–16, 1989

  36. [44]

    Fergus I. M. Craik and Robert S. Lockhart. Levels of processing: A framework for memory research. Journal of Verbal Learning and Verbal Behavior , 11(6):671– 684, 1972

  37. [45]

    Artificial intelligence in higher education: The state of the field

    Helen Crompton and Diane Burke. Artificial intelligence in higher education: The state of the field. International Journal of Educational Technology in Higher Education, 20(1):22, 2023

  38. [46]

    The friendly orange glow: The untold story of the PLATO system and the dawn of cyberculture

    Brian Dear. The friendly orange glow: The untold story of the PLATO system and the dawn of cyberculture. Pantheon Books, 2017

  39. [47]

    BERT: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language T...

  40. [48]

    Dillahunt, Zengguang Wang, and Stephanie D

    Tawanna R. Dillahunt, Zengguang Wang, and Stephanie D. Teasley. Democra- tizing higher education: Exploring MOOC use among those who cannot afford a formal education. The International Review of Research in Open and Distributed Learning, 15(5):177–196, 2014

  41. [49]

    Video browse by direct manipulation

    Pierre Dragicevic, Gonzalo Ramos, Jacobo Bibliowitcz, Derek Nowrouzezahrai, Ravin Balakrishnan, and Karan Singh. Video browse by direct manipulation. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 237–246, 2008

  42. [50]

    Drucker, Asta Glatzer, Steven De Mar, and Curtis Wong

    Steven M. Drucker, Asta Glatzer, Steven De Mar, and Curtis Wong. SmartSkip: Consumer level browse and skipping of digital video content. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 219–226, 2002

  43. [51]

    Why college students watch streaming drama at higher playback speed: The uses and gratifications perspective

    Songshuang Duan and Xiaoqian Chen. Why college students watch streaming drama at higher playback speed: The uses and gratifications perspective. In Pro- ceedings of the International Joint Conference on Information, Media and Engi- neering, 2019

  44. [52]

    Anders Ericsson, Ralf T

    K. Anders Ericsson, Ralf T. Krampe, and Clemens Tesch-Römer. The role of deliberate practice in the acquisition of expert performance.Psychological Review, 100(3):363–406, 1993. 93

  45. [53]

    Metacognition and self-regulation

    Evidence for Learning. Metacognition and self-regulation. Teaching and Learning Toolkit, 2021. Review last updated July 2021

  46. [54]

    Investigating the effects of artificial intelligence-assisted language learning strategies on cognitive load and learning outcomes: A comparative study

    Lijuan Feng. Investigating the effects of artificial intelligence-assisted language learning strategies on cognitive load and learning outcomes: A comparative study. Journal of Educational Computing Research , 62(8):1741–1774, 2025. First pub- lished online 31 August 2024; iss...

  47. [55]

    Effects of experience on non-native speakers’ production and perception of English vowels

    James Emil Flege, Ocke-Schwen Bohn, and Sunyoung Jang. Effects of experience on non-native speakers’ production and perception of English vowels. Journal of Phonetics, 25(4):437–470, 1997

  48. [56]

    Au- tomatic speech recognition predicts speech intelligibility and comprehension for listeners with simulated age-related hearing loss

    Lionel Fontan, Isabelle Ferrané, Jérôme Farinas, Julien Pinquier, Julien Tardieu, Cynthia Magnen, Pascal Gaillard, Xavier Aumont, and Christian Füllgrabe. Au- tomatic speech recognition predicts speech intelligibility and comprehension for listeners with simulated age-related ...

  49. [57]

    Adam Fouse, Nadir Weibel, Edwin Hutchins, and James D. Hollan. ChronoViz: A system for supporting navigation of time-coded data. In Proceedings of the CHI Conference on Human Factors in Computing Systems Extended Abstracts , pages 299–304, 2011

  50. [58]

    Pedagogy of the oppressed

    Paulo Freire. Pedagogy of the oppressed. Seabury Press, 1970

  51. [59]

    The perceptual learning of time- compressed speech: A comparison of training protocols with different levels of difficulty

    Y afit Gabay, Avi Karni, and Karen Banai. The perceptual learning of time- compressed speech: A comparison of training protocols with different levels of difficulty. PLoS ONE, 12(5):e0176488, 2017

  52. [60]

    J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren. DARPA TIMIT acoustic phonetic continuous speech corpus CDROM. Technical Report NISTIR-4930, National Institute of Standards and Technology (NIST), 1993

  53. [61]

    Giegerich

    Heinz J. Giegerich. English phonology: An introduction . Cambridge Textbooks in Linguistics. Cambridge University Press, 1992

  54. [62]

    The complete HyperCard handbook

    Danny Goodman. The complete HyperCard handbook. Bantam Books, 1987

  55. [63]

    Graves, S

    A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural nets. In Proceedings of the International Conference on Machine Learning , 2006

  56. [64]

    Speech recognition with deep recurrent neural networks

    Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. Speech recognition with deep recurrent neural networks. In Proceedings of the International Confer- ence on Acoustics, Speech and Signal Processing , pages 6645–6649, 2013

  57. [65]

    A survey on self-supervised learning: Algorithms, applications, and fu- ture trends

    Jie Gui, Tuo Chen, Jing Zhang, Qiong Cao, Zhenan Sun, Hao Luo, and Dacheng Tao. A survey on self-supervised learning: Algorithms, applications, and fu- ture trends. IEEE Transactions on Pattern Analysis and Machine Intelligence , 46(12):9052–9071, 2024

  58. [66]

    Guo, Juho Kim, and Rob Rubin

    Philip J. Guo, Juho Kim, and Rob Rubin. How video production affects student engagement: An empirical study of MOOC videos. In Proceedings of the ACM Conference on Learning at Scale , pages 41–50. ACM, 2014. 94

  59. [67]

    M. P . J. Habgood and S. E. Ainsworth. Motivating children to learn effectively: Exploring the value of intrinsic integration in educational games. The Journal of the Learning Sciences , 20(2):169–206, 2011

  60. [68]

    Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, and An- drew Y . Ng. Deep speech: Scaling up end-to-end speech recognition. arXiv, 2014

  61. [69]

    Exploring collaborative decision- making: A quasi-experimental study of human and generative AI interaction.Tech- nology in Society, 78:102662, 2024

    Xinyue Hao, Emrah Demir, and Daniel Eyers. Exploring collaborative decision- making: A quasi-experimental study of human and generative AI interaction.Tech- nology in Society, 78:102662, 2024

  62. [70]

    Hardison

    Debra M. Hardison. Generalization of computer-assisted prosody training: Quan- titative and qualitative findings. Language Learning & Technology , 8(1):34–52, 2004

  63. [71]

    Hartshorne, Joshua B

    Joshua K. Hartshorne, Joshua B. Tenenbaum, and Steven Pinker. A critical pe- riod for second language acquisition: Evidence from 2/3 million English speakers. Cognition, 177:263–277, 2018

  64. [72]

    Teaching machines to read and com- prehend

    Karl Moritz Hermann, Tomáš Kočiský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. Teaching machines to read and com- prehend. In Proceedings of the International Conference on Neural Information Processing Systems, pages 1693–1701, 2015

  65. [73]

    Students’ and instructors’ use of mas- sive open online courses (MOOCs): Motivations and challenges

    Khe Foon Hew and Wing Sum Cheung. Students’ and instructors’ use of mas- sive open online courses (MOOCs): Motivations and challenges. Educational Research Review, 12:45–58, 2014

  66. [74]

    EgoScanning: Quickly scanning first-person videos with egocentric elastic timelines

    Keita Higuchi, Ryo Y onetani, and Y oichi Sato. EgoScanning: Quickly scanning first-person videos with egocentric elastic timelines. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 6536–6546. Associ- ation for Computing Machinery, 2017

  67. [75]

    Deep metric learning using triplet network

    Elad Hoffer and Nir Ailon. Deep metric learning using triplet network. Similarity- Based Pattern Recognition, pages 84–92, 2015

  68. [76]

    Designing for human-AI complementarity in K-12 education

    Kenneth Holstein and Vincent Aleven. Designing for human-AI complementarity in K-12 education. AI Magazine, 43(2):239–248, 2022

  69. [77]

    The struggle for recognition: The moral grammar of social con- flicts

    Axel Honneth. The struggle for recognition: The moral grammar of social con- flicts. MIT Press, 1995

  70. [78]

    Hershey, Tim K

    Chiori Hori, Takaaki Hori, Teng- Y ok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi. Attention-based multimodal fusion for video description. In Proceedings of the IEEE International Conference on Computer Vision, pages 4193–4202, 2017

  71. [79]

    Principles of mixed-initiative user interfaces

    Eric Horvitz. Principles of mixed-initiative user interfaces. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 159–166, 1999

  72. [80]

    HuBERT: Self-supervised speech representation learning by masked prediction of hidden units

    Wei-Ning Hsu, Benjamin Bolte, Y ao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units. IEEE/ACM Trans- actions on Audio, Speech and Language Processing , ...

  73. [81]

    Cognition in the Wild

    Edwin Hutchins. Cognition in the Wild . MIT Press, 1995. 95

  74. [82]

    Deschooling society

    Ivan Illich. Deschooling society. Harper & Row, 1971

  75. [83]

    Designing participatory AI: Cre- ative professionals`worries and expectations about generative AI

    Nanna Inie, Jeanette Falk, and Steve Tanimoto. Designing participatory AI: Cre- ative professionals`worries and expectations about generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems Extended Ab- stracts, pages 1–8, 2023

  76. [84]

    Tangible bits: Towards seamless interfaces be- tween people, bits and atoms

    Hiroshi Ishii and Brygg Ullmer. Tangible bits: Towards seamless interfaces be- tween people, bits and atoms. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 234–241, 1997

  77. [85]

    Word perception in fast speech: Artificially time-compressed vs

    Esther Janse. Word perception in fast speech: Artificially time-compressed vs. naturally produced fast speech. Speech Communication, 42:155–173, 2004

  78. [86]

    Speech recognition performance as an ef- fective perceived quality predictor

    Wenyu Jiang and Henning Schulzrinne. Speech recognition performance as an ef- fective perceived quality predictor. InProceedings of the Tenth IEEE International Workshop on Quality of Service , pages 269–275. IEEE, 2002

  79. [87]

    Us- ing ChatGPT for course curriculum design: A systematic review

    Michelle Celine J”orgens, Florian Beier, Sebastian Kreibich, and Dirk Werth. Us- ing ChatGPT for course curriculum design: A systematic review. In The Paris Conference on Education 2024: Official Conference Proceedings, pages 549–561, 2024

  80. [88]

    A survey of speaker recognition: Fundamental theories, recognition methods and opportunities

    Muhammad Mohsin Kabir, Muhammad Firoz Mridha, Jungpil Shin, Israt Jahan, and Abu Quwsar Ohi. A survey of speaker recognition: Fundamental theories, recognition methods and opportunities. IEEE Access, 9:79236–79263, 2021

  81. [89]

    The expertise rever- sal effect

    Slava Kalyuga, Paul Ayres, Paul Chandler, and John Sweller. The expertise rever- sal effect. Educational Psychologist, 38(1):23–31, 2003

  82. [90]

    Educational scalability in MOOCs: Analysing instructional designs to find best practices

    Julia Kasch, Peter Van Rosmalen, and Marco Kalz. Educational scalability in MOOCs: Analysing instructional designs to find best practices. Computers & Education, 161:104054, 2021

  83. [91]

    ChatGPT for good? on opportunities and challenges of large language models for education

    Enkelejda Kasneci, Kathrin Seßler, Stefan Küchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann, Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Saile...

  84. [92]

    Efficient video viewing system for racquet sports with automatic summarization focusing on rally scenes

    Shunya Kawamura, Tsukasa Fukusato, Tatsunori Hirai, and Shigeo Morishima. Efficient video viewing system for racquet sports with automatic summarization focusing on rally scenes. In Proceedings of the ACM SIGGRAPH Posters, 2014

  85. [93]

    Robin H. Kay. Exploring the use of video podcasts in education: A comprehensive review of the literature. Computers in Human Behavior , 28(3):820–831, 2012

  86. [94]

    Explainable artificial intelligence in education: A com- prehensive review

    Hassan Khosravi, Simon Buckingham Shum, Guanliang Chen, Cristina Conati, Yi- Shan Tsai, Judy Kay, Simon Knight, Roberto Martinez-Maldonado, Shazia Sadiq, and Dragan Gašević. Explainable artificial intelligence in education: A com- prehensive review. Computers and Education: Ar...

  87. [95]

    Digital game-based learning: Towards an experiential gaming model

    Kristian Kiili. Digital game-based learning: Towards an experiential gaming model. The Internet and Higher Education , 8(1):13–24, 2005. 96

  88. [96]

    Automatic pronunciation assessment using self-supervised speech representation learning

    Eesung Kim, Jae-Jin Jeon, Hyeji Seo, and Hoon Kim. Automatic pronunciation assessment using self-supervised speech representation learning. In Proceedings of INTERSPEECH, pages 1411–1415. ISCA, 2022

  89. [97]

    Generic speech summarization of transcribed lecture videos: Using tags and their semantic relations

    Hyun Hee Kim and Y ong Ho Kim. Generic speech summarization of transcribed lecture videos: Using tags and their semantic relations. Journal of the Association for Information Science and Technology, 67(2):366–379, 2016

  90. [98]

    Guo, Daniel T

    Juho Kim, Philip J. Guo, Daniel T. Seaton, Piotr Mitros, Krzysztof Z. Gajos, and Robert C. Miller. Understanding in-video dropouts and interaction peaks in online lecture videos. In Proceedings of the First ACM Conference on Learning at Scale, pages 31–40. ACM, 2014

  91. [99]

    Kingma and Jimmy Ba

    Diederik P . Kingma and Jimmy Ba. Adam: A method for stochastic optimiza- tion. In Proceedings of the International Conference on Learning Representations, 2015

  92. [100]

    David A. Kolb. Experiential Learning: Experience as the Source of Learning and Development. Prentice-Hall, Englewood Cliffs, NJ, 1984

  93. [101]

    Kulik and J

    James A. Kulik and J. D. Fletcher. Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1):42–78, 2016

  94. [102]

    CinemaGazer: A system for watching video at very high speed

    Kazutaka Kurihara. CinemaGazer: A system for watching video at very high speed. In Proceedings of the Workshop on Advanced Visual Interfaces , 2012

  95. [103]

    Is faster better? a study of video playback speed

    David Lang, Guanliang Chen, Kathy Mirzaei, and Andreas Paepcke. Is faster better? a study of video playback speed. In Proceedings of the International Conference on Learning Analytics & Knowledge, pages 260–269. Association for Computing Machinery, 2020

  96. [104]

    Situated learning: Legitimate peripheral partici- pation

    Jean Lave and Etienne Wenger. Situated learning: Legitimate peripheral partici- pation. Cambridge University Press, 1991

  97. [105]

    Berg, and Mohit Bansal

    Jie Lei, Tamara L. Berg, and Mohit Bansal. Detecting moments and highlights in videos via natural language queries. InAdvances in Neural Information Processing Systems, volume 34, pages 11846–11858, 2021

  98. [106]

    Chen, and Chin-Hui Lee

    Wei Li, Kehuang Li, Sabato Marco Siniscalchi, Nancy F. Chen, and Chin-Hui Lee. Detecting mispronunciations of L2 learners and providing corrective feed- back using knowledge-guided and data-driven decision trees. In Proceedings of INTERSPEECH, 2016

  99. [107]

    Enhancing length generalization for attention based knowledge tracing models with linear biases

    Xueyi Li, Y ouheng Bai, Teng Guo, Zitao Liu, Y aying Huang, Xiangyu Zhao, Feng Xia, Weiqi Luo, and Jian Weng. Enhancing length generalization for attention based knowledge tracing models with linear biases. In Proceedings of the Thirty- Third International Joint Conference on ...

  100. [108]

    Focal loss for dense object detection

    Tsung- Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, pages 2980–2988, 2017. 97

  101. [109]

    Audio self-supervised learn- ing: A survey

    Shuo Liu, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Kun Qian, Xin Jing, Alexander Kathan, Bin Hu, and Bjoern W Schuller. Audio self-supervised learn- ing: A survey. Patterns, 3(12):100616, 2022

  102. [110]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Pro- ceedings of the 7th International Conference on Learning Representations , 2019

  103. [111]

    The modality principle in multimedia learning

    Renae Low and John Sweller. The modality principle in multimedia learning. In Richard E. Mayer, editor,The Cambridge Handbook of Multimedia Learning (2nd ed.), pages 227–246. Cambridge University Press, 2014

  104. [112]

    Automatic speech recognition: A survey

    Mishaim Malik, Muhammad Malik, Khawar Mehmood, and Imran Makhdoom. Automatic speech recognition: A survey. Multimedia Tools and Applications , 80:9411–9457, 2021

  105. [113]

    Florence Martin, Ting Sun, and Carl D. Westine. A systematic review of research on online teaching and learning from 2009 to 2018. Computers & Education , 159:104009, 2020

  106. [114]

    Individual differences in working memory capacity moderate effects of post-learning activity on memory consolidation over the long term

    Markus Martini, Robert Marhenke, Caroline Martini, Sonja Rossi, and Pierre Sachse. Individual differences in working memory capacity moderate effects of post-learning activity on memory consolidation over the long term. Scientific Re- ports, 10(1):17976, 2020

  107. [115]

    Swift: Reducing the effects of latency in online video scrubbing

    Justin Matejka, Tovi Grossman, and George Fitzmaurice. Swift: Reducing the effects of latency in online video scrubbing. In Proceedings of the CHI Confer- ence on Human Factors in Computing Systems , pages 637–646. Association for Computing Machinery, 2012

  108. [116]

    Mayer, editor

    Richard E. Mayer, editor. The Cambridge handbook of multimedia learning. Cam- bridge University Press, 2005

  109. [117]

    Richard E. Mayer. Cognitive theory of multimedia learning. In Richard E. Mayer, editor, The Cambridge handbook of multimedia learning, pages 31–48. Cambridge University Press, 2005

  110. [118]

    Richard E. Mayer. Multimedia learning . Cambridge University Press, 3rd ed. edition, 2020

  111. [119]

    Mayer, Kristina Sobko, and Patricia D

    Richard E. Mayer, Kristina Sobko, and Patricia D. Mautone. Social cues in mul- timedia learning: Role of speaker’s voice. Journal of Educational Psychology , 95(2):419–425, 2003

  112. [120]

    Montreal Forced Aligner: Trainable text-speech alignment us- ing Kaldi

    Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Mor- gan Sonderegger. Montreal Forced Aligner: Trainable text-speech alignment us- ing Kaldi. In Proceedings of INTERSPEECH, pages 498–502, 2017

  113. [121]

    Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto

    Brian McFee, Colin Raffel, Dawen Liang, Daniel P . Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. Librosa: Audio and music signal analysis in Python. In Proceedings of the Python in Science Conference , pages 18–24, 2015

  114. [122]

    Development of En- glish speech database spoken by Japanese learners

    Nobuaki Minematsu, Y oshihiro Tomiyama, Kei Y oshimoto, Katsumasa Shimizu, Seiichi Nakagawa, Masatake Dantsuji, and Shozo Makino. Development of En- glish speech database spoken by Japanese learners. In Proceedings of the CO- COSDA Workshop 2001, pages 76–81, 2001. 98

  115. [123]

    English speech database read by Japanese learners for CALL system development

    Nobuaki Minematsu, Y oshihiro Tomiyama, Kei Y oshimoto, Katsumasa Shimizu, Seiichi Nakagawa, Masatake Dantsuji, and Shozo Makino. English speech database read by Japanese learners for CALL system development. InProceedings of the International Conference on Language Resources ...

  116. [124]

    Punya Mishra and Matthew J. Koehler. Technological pedagogical content knowledge: A framework for teacher knowledge. Teachers College Record , 108(6):1017–1054, 2006

  117. [125]

    Roxana Moreno and Richard E. Mayer. Cognitive principles of multimedia learn- ing: The role of modality and contiguity. Journal of Educational Psychology , 91(2):358–368, 1999

  118. [126]

    Multimodal data fusion in learning analyt- ics: A systematic review

    Su Mu, Meng Cui, and Xiaodi Huang. Multimodal data fusion in learning analyt- ics: A systematic review. Sensors, 20(23):6856, 2020

  119. [127]

    Murphy, Kara M

    Dillon H. Murphy, Kara M. Hoover, Karina Agadzhanyan, Jesse C. Kuehn, and Alan D. Castel. Learning in double time: The effect of lecture video speed on immediate and delayed comprehension. Applied Cognitive Psychology, 36(1):69– 82, 2022

  120. [128]

    The pedagogy– technology interface in computer assisted pronunciation training

    Ambra Neri, Catia Cucchiarini, Helmer Strik, and Lou Boves. The pedagogy– technology interface in computer assisted pronunciation training. Computer As- sisted Language Learning, 15(5):441–467, 2002

  121. [129]

    The effective- ness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis

    Thuy Thi-Nhu Ngo, Howard Hao-Jan Chen, and Kyle Kuo-Wei Lai. The effective- ness of automatic speech recognition in ESL/EFL pronunciation: A meta-analysis. ReCALL, 36(1):4–21, 2024

  122. [130]

    Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran

    Jeremy Nixon, Michael W. Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran. Measuring calibration in deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , pages 38–41, 2019

  123. [131]

    Donald A. Norman. The design of everyday things: Revised and expanded edition . Basic Books, 2013

  124. [132]

    GPT-4V(ision) system card

    OpenAI. GPT-4V(ision) system card. Technical report, OpenAI, 2023

  125. [133]

    Pacheco, Charley W

    Matheus M. Pacheco, Charley W. Lafe, and Karl M. Newell. Search strategies in the perceptual-motor workspace and the acquisition of coordination, control, and skill. Frontiers in Psychology, 10:1874, 2019

  126. [134]

    The vowel game: Continuous real- time visualization for pronunciation learning with vowel charts

    Annu Paganus, Vesa-Petteri Mikkonen, Tomi Mäntylä, Sami Nuuttila, Jouni Isoaho, Olli Aaltonen, and Tapio Salakoski. The vowel game: Continuous real- time visualization for pronunciation learning with vowel charts. In Advances in Natural Language Processing, pages 696–703. Spri...

  127. [135]

    Mul- timodal abstractive summarization for How2 videos

    Shruti Palaskar, Jindřich Libovický, Spandana Gella, and Florian Metze. Mul- timodal abstractive summarization for How2 videos. In Anna Korhonen, David Traum, and Lluís Màrquez, editors, Proceedings of the Annual Meeting of the As- sociation for Computational Linguistics , pag...

  128. [136]

    Lib- rispeech: An ASR corpus based on public domain audio books

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Lib- rispeech: An ASR corpus based on public domain audio books. In Proceedings of 99 the IEEE International Conference on Acoustics, Speech and Signal Processing , pages 5206–5210, 2015

  129. [137]

    Mindstorms: Children, computers, and powerful ideas

    Seymour Papert. Mindstorms: Children, computers, and powerful ideas . Basic Books, 1980

  130. [138]

    Pastore and Albert D

    Raymond S. Pastore and Albert D. Ritzhaupt. Using time-compression to make multimedia learning more efficient: Current research and practice. TechTrends, 59:66–74, 2015

  131. [139]

    PyTorch: An imperative style, high-performance deep learning library

    Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gre- gory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Y ang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...

  132. [140]

    Video di- gests: A browsable, skimmable format for informational lecture videos

    Amy Pavel, Colorado Reed, Björn Hartmann, and Maneesh Agrawala. Video di- gests: A browsable, skimmable format for informational lecture videos. In Pro- ceedings of the 27th Annual ACM Symposium on User Interface Software and Technology, pages 573–582. Association for Computin...

  133. [141]

    Tuned models of peer assessment in MOOCs

    Chris Piech, Jonathan Huang, Zhenghao Chen, Chuong Do, Andrew Ng, and Daphne Koller. Tuned models of peer assessment in MOOCs. In Proceedings of the Educational Data Mining , pages 153–160, 2013

  134. [142]

    LLMs in education: Evaluation GPT and BERT models in student comment classification

    Anabel Pilicita and Enrique Barra. LLMs in education: Evaluation GPT and BERT models in student comment classification. Multimodal Technologies and Interaction, 9(5):44, 2025

  135. [143]

    Plass, Bruce D

    Jan L. Plass, Bruce D. Homer, and Charles K. Kinzer. Foundations of game-based learning. Educational Psychologist, 50(4):258–283, 2015

  136. [144]

    The tacit dimension

    Michael Polanyi. The tacit dimension . Doubleday, 1966

  137. [145]

    Content- aware dynamic timeline for video browse

    Suporn Pongnumkul, Jue Wang, Gonzalo Ramos, and Michael Cohen. Content- aware dynamic timeline for video browse. In Proceedings of the annual ACM symposium on User interface software and technology , pages 139–142, 2010

  138. [146]

    Discriminatively trained acoustic models for improving mispronunciation detection and diagnosis in computer aided pronunciation training (CAPT)

    Xiaojun Qian, Frank Soong, and Helen Meng. Discriminatively trained acoustic models for improving mispronunciation detection and diagnosis in computer aided pronunciation training (CAPT). In Proceedings of INTERSPEECH, 2010

  139. [147]

    Robust speech recognition via large-scale weak supervision

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, p...

  140. [148]

    Improv- ing language understanding by generative pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improv- ing language understanding by generative pre-training. Technical report, OpenAI, 2018

  141. [149]

    Language models are unsupervised multitask learners

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. Technical Re- port TR-2019-1, OpenAI, 2019. 100

  142. [150]

    Radwan, Nancy M

    Nisreen I. Radwan, Nancy M. Salem, and Mohamed I. El Adawy. Histogram correlation for video scene change detection. In Proceedings of the Second Inter- national Conference on Computer Science, Engineering and Applications , pages 765–773, 2012

  143. [151]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Y anqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67, 2020

  144. [152]

    Y amamoto Ravenor

    R. Y amamoto Ravenor. On the adequacy of Elsa Speak in formal education: A survey of teacher-users. Advances in Artificial Intelligence and Machine Learning, 04:2387–2394, 2024

  145. [153]

    A theory of justice

    John Rawls. A theory of justice . Harvard University Press, 1971

  146. [154]

    ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope

    Partha Pratim Ray. ChatGPT: A comprehensive review on background, applica- tions, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3:121–154, 2023

  147. [155]

    Brian J. Reiser. Scaffolding complex learning: The mechanisms of structuring and problematizing student work. The Journal of the Learning Sciences , 13(3):273– 304, 2004

  148. [156]

    Faster R-CNN: To- wards real-time object detection with region proposal networks

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster R-CNN: To- wards real-time object detection with region proposal networks. In C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, editors,Proceedings of the Ad- vances in Neural Information Processing Syste...

  149. [157]

    Scratch: Programming for all

    Mitchel Resnick, John Maloney, Andrés Monroy-Hernández, Natalie Rusk, Eve- lyn Eastmond, Karen Brennan, Amon Millner, Eric Rosenbaum, Jay Silver, Brian Silverman, and Y asmin Kafai. Scratch: Programming for all. Communications of the ACM, 52(11):60–67, 2009

  150. [158]

    Risko, Nicola Anderson, Amrita Sarwal, Megan Engelhardt, and Alan Kingstone

    Evan F. Risko, Nicola Anderson, Amrita Sarwal, Megan Engelhardt, and Alan Kingstone. Everyday attention: Variation in mind wandering and memory in a lecture. Applied Cognitive Psychology, 26(2):234–242, 2012

  151. [159]

    Rogerson-Revell

    Pamela M. Rogerson-Revell. Computer-assisted pronunciation training (CAPT): Current issues and future directions. RELC Journal, 52(1):189–205, 2021

  152. [160]

    Tradition or innovation: A comparison of modern ASR methods for forced alignment

    Rotem Rousso, Eyal Cohen, Joseph Keshet, and Eleanor Chodroff. Tradition or innovation: A comparison of modern ASR methods for forced alignment. In Pro- ceedings of INTERSPEECH 2024 , pages 1525–1529, 2024

  153. [161]

    Chat- GPT in lesson preparation: A teacher choices trial

    Palak Roy, Helen Poet, Ruth Staunton, Katherine Aston, and David Thomas. Chat- GPT in lesson preparation: A teacher choices trial. Technical report, National Foundation for Educational Research, December 2024

  154. [162]

    Effects of second language pronunciation teach- ing revisited: A proposed measurement framework and meta-analysis

    Kazuya Saito and Luke Plonsky. Effects of second language pronunciation teach- ing revisited: A proposed measurement framework and meta-analysis. Language Learning, 69(3):652–708, 2019

  155. [163]

    Scheirer and M

    E. Scheirer and M. Slaney. Construction and evaluation of a robust multifeature speech/music discriminator. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , volume 2, pages 1331–1334, 1997. 101

  156. [164]

    wav2vec: Unsupervised pre-training for speech recognition.Proceedings of INTERSPEECH, pages 3465–3469, 2019

    Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised pre-training for speech recognition.Proceedings of INTERSPEECH, pages 3465–3469, 2019

  157. [165]

    Generative AI in education: Past, present, and future

    Tony Sheehan. Generative AI in education: Past, present, and future. EDUCAUSE Review, 2023

  158. [166]

    Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Y ang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R

    Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Y ang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R. J. Skerrv-Ryan, Rif A. Saurous, Y annis Agiomyrgiannakis, and Y onghui Wu. Natural TTS syn- thesis by conditioning Wavenet on MEL spectrogram predi...

  159. [167]

    Sherwood

    Bruce A. Sherwood. The computer speaks. IEEE Spectrum, 16(8):18–25, 1979

  160. [168]

    Direct manipulation: A step beyond programming languages

    Ben Shneiderman. Direct manipulation: A step beyond programming languages. Computer, 16(8):57–69, 1983

  161. [169]

    Direct manipulation for comprehensible, predictable and con- trollable user interfaces

    Ben Shneiderman. Direct manipulation for comprehensible, predictable and con- trollable user interfaces. In Proceedings of the International Conference on Intel- ligent User Interfaces, pages 33–39, 1997

  162. [170]

    Human-centered AI

    Ben Shneiderman. Human-centered AI. Oxford University Press, 2022

  163. [171]

    Direct manipulation versus interface agents

    Ben Shneiderman and Pattie Maes. Direct manipulation versus interface agents. Interactions, 4(6):42–61, 1997

  164. [172]

    Gamification in mobile-assisted language learning: A sys- tematic review of Duolingo literature from public release of 2012 to early 2020

    Mitchell Shortt, Shantanu Tilak, Irina Kuznetcova, Bethany Martens, and Ba- batunde Akinkuolie. Gamification in mobile-assisted language learning: A sys- tematic review of Duolingo literature from public release of 2012 to early 2020. Computer Assisted Language Learning, 36(3)...

  165. [173]

    Augmented visual, auditory, haptic and multimodal feedback in motor learning: A review

    Roland Sigrist, Georg Rauter, Robert Riener, and Peter Wolf. Augmented visual, auditory, haptic and multimodal feedback in motor learning: A review. Psycho- nomic Bulletin & Review , 20:21–53, 2013

  166. [174]

    Herbert A. Simon. Designing organizations for an information-rich world. In Martin Greenberger, editor, Computers, Communications, and the Public Interest, pages 37–72. Johns Hopkins Press, 1971

  167. [175]

    The critical period hypothesis for L2 acquisition: An unfalsifiable embarrassment? Languages, 6(3):149, 2021

    David Singleton and Justyna Leśniewska. The critical period hypothesis for L2 acquisition: An unfalsifiable embarrassment? Languages, 6(3):149, 2021

  168. [176]

    R. Smith. An overview of the Tesseract OCR Engine. In Proceedings of the International Conference on Document Analysis and Recognition, volume 2, pages 629–633, 2007

  169. [177]

    Dropout: A simple way to prevent neural networks from overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research , 15(56):1929–1958, 2014

  170. [178]

    Sullivan, Shailesh S

    Katherine J. Sullivan, Shailesh S. Kantak, and Patricia A. Burtner. Motor learning in children: Feedback effects on skill acquisition. Physical Therapy, 88(6):720– 732, 2008. 102

  171. [179]

    Sutherland

    Ivan E. Sutherland. Sketchpad: a man-machine graphical communication system. In Proceedings of the May 21-23, 1963, Spring Joint Computer Conference, pages 329–346. ACM, 1963

  172. [180]

    Cognitive load during problem solving: Effects on learning

    John Sweller. Cognitive load during problem solving: Effects on learning. Cog- nitive Science, 12(2):257–285, 1988

  173. [181]

    Cognitive load theory

    John Sweller. Cognitive load theory. In Psychology of learning and motivation , volume 55, pages 37–76. Elsevier, 2011

  174. [182]

    John Sweller, Jeroen J. G. van Merriënboer, and Fred Paas. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review, 31:261– 292, 2019

  175. [183]

    The politics of recognition

    Charles Taylor. The politics of recognition. In Amy Gutmann, editor, Multicul- turalism, pages 25–74. Princeton University Press, 1994

  176. [184]

    Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers

    Cristian Tejedor-García, Valentín Cardeñoso Payo, and David Escudero-Mancebo. Automatic speech recognition (ASR) systems applied to pronunciation assessment of L2 Spanish for Japanese speakers. Applied Sciences, 11(15), 2021

  177. [185]

    Theepan Tharumalingam, Brady R. T. Roberts, Jonathan M. Fawcett, and Evan F. Risko. Increasing video lecture playback speed can impair test performance—a meta-analysis. Educational Psychology Review, 37, 2025

  178. [186]

    How Khan Academy is changing the rules of education

    Clive Thompson. How Khan Academy is changing the rules of education. Wired, July 2011

  179. [187]

    Ron I. Thomson. Measurement of accentedness, intelligibility and comprehensi- bility. In Okim Kang and April Ginther, editors, Assessment in second language pronunciation, pages 11–29. Routledge, 2018

  180. [188]

    EduQate: Generating adap- tive curricula through RMABs in education settings

    Sidney Tio, Dexun Li, and Pradeep Varakantham. EduQate: Generating adap- tive curricula through RMABs in education settings. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems , pages 2042–2050, 2025. ACM Digital Library identifier ...

  181. [189]

    Video abstraction: A systematic review and classification

    Ba Tu Truong and Svetha Venkatesh. Video abstraction: A systematic review and classification. ACM Transactions on Multimedia Computing, Communications, and Applications, 3(1), 2007

  182. [190]

    WaveNet: A generative model for raw audio

    Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. InProceedings of the ISCA Workshop on Speech Synthesis Workshop, page 125, 2016

  183. [191]

    Audio summarization for podcasts

    Aneesh Vartakavi, Amanmeet Garg, and Zafar Rafii. Audio summarization for podcasts. In Proceedings of the European Signal Processing Conference , pages 431–435. IEEE, 2021

  184. [192]

    Gomez, Ł ukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems , vol- ume 30, pages 5998–6008, 2017. 103

  185. [193]

    Vygotsky

    Lev S. Vygotsky. Mind in society: The development of higher psychological pro- cesses. Harvard University Press, 1978

  186. [194]

    Web Audio API

    W3C Audio Working Group. Web Audio API. W3C Recommendation, 2021. 17 June 2021

  187. [195]

    WebRTC 1.0: Real-time communication between browsers

    W3C Web Real-Time Communications Working Group. WebRTC 1.0: Real-time communication between browsers. W3C Recommendation, 2021. 26 January 2021

  188. [196]

    Fairseq S2T: Fast speech-to-text modeling with Fairseq

    Changhan Wang, Yun Tang, Xutai Ma, Anne Wu, Dmytro Okhonko, and Juan Pino. Fairseq S2T: Fast speech-to-text modeling with Fairseq. In Proceedings of the Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Con...

  189. [197]

    Artificial-intelligence-generated content with diffusion models: A literature review

    Xiaolong Wang, Zhijian He, and Xiaojiang Peng. Artificial-intelligence-generated content with diffusion models: A literature review. Mathematics, 12(7):977, 2024

  190. [198]

    Yuxuan Wang, R. J. Skerry-Ryan, Daisy Stanton, Y onghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Y ang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc V . Le, Y annis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous. Tacotron: Towards end-to-end speech synthesis. In Proceedings o...

  191. [199]

    Six views of embodied cognition

    Margaret Wilson. Six views of embodied cognition. Psychonomic Bulletin & Review, 9:625–636, 2002

  192. [200]

    Witt and Steve J

    Silke M. Witt and Steve J. Y oung. Phone-level pronunciation scoring and assess- ment for interactive language learning. Speech Communication , 30(2):95–108, 2000

  193. [201]

    Bruner, and Gail Ross

    David Wood, Jerome S. Bruner, and Gail Ross. The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry , 17(2):89–100, 1976

  194. [202]

    David R. Woolley. PLATO: The emergence of on-line community. Computer- Mediated Communication Magazine, 1(3):5, 1994

  195. [203]

    Improving student-AI interaction through ped- agogical prompting: An example in computer science education

    Ruiwei Xiao, Xinying Hou, Runlong Y e, Majeed Kazemitabaar, Nicholas Diana, Michael Liut, and John Stamper. Improving student-AI interaction through ped- agogical prompting: An example in computer science education. arXiv preprint arXiv:2506.19107v1, 2025. Submitted 23 June 2025

  196. [204]

    Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig

    Wayne Xiong, Jasha Droppo, Xuedong Huang, Frank Seide, Michael L. Seltzer, Andreas Stolcke, Dong Yu, and Geoffrey Zweig. Toward human parity in con- versational speech recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, pages 2410–2423, 2017

  197. [205]

    Conformer-based speech recognition on extreme edge-computing devices

    Mingbin Xu, Alex Jin, Sicheng Wang, Mu Su, Tim Ng, Henry Mason, Shiyi Han, Zhihong Lei, Y aqiao Deng, Zhen Huang, and Mahesh Krishnamoorthy. Conformer-based speech recognition on extreme edge-computing devices. In Pro- ceedings of the Conference of the North American Chapter o...

  198. [206]

    VidChapters-7M: Video chapters at scale

    Antoine Y ang, Arsha Nagrani, Ivan Laptev, Josef Sivic, and Cordelia Schmid. VidChapters-7M: Video chapters at scale. arXiv, 2023. arXiv:2309.13952. NeurIPS 2023 Datasets and Benchmarks Track

  199. [207]

    Re-examining whether, why, and how human-AI interaction is uniquely difficult to design

    Qian Y ang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman. Re-examining whether, why, and how human-AI interaction is uniquely difficult to design. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1–13, 2020

  200. [208]

    SoftVideo: Improving the learning experience of software tutorial videos with collective interaction data

    Saelyne Y ang, Jisu Yim, Aitolkyn Baigutanova, Seoyoung Kim, Minsuk Chang, and Juho Kim. SoftVideo: Improving the learning experience of software tutorial videos with collective interaction data. In Proceedings of the International Con- ference on Intelligent User Interfaces, ...

  201. [209]

    Lin, Andy T

    Shu-wen Y ang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakho- tia, Yist Y . Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu- Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahm...

  202. [210]

    Personalized video summarization based on behavior of viewer

    Atsuo Y oshitaka and Kazuya Sawada. Personalized video summarization based on behavior of viewer. In Proceedings of the International Conference on Signal Image Technology and Internet Based Systems, pages 661–667, 2012

  203. [211]

    Predictive video analytics in online courses: A systematic literature review.Technology, Knowledge and Learning, 29:1907–1937, 2024

    Ozan Raşit Yürüm, Tuğba Taşkaya-Temizel, and Soner Yıldırım. Predictive video analytics in online courses: A systematic literature review.Technology, Knowledge and Learning, 29:1907–1937, 2024

  204. [212]

    Zekveld, Sophia E

    Adriana A. Zekveld, Sophia E. Kramer, and Joost M. Festen. Cognitive load during speech perception in noise: The influence of age, hearing loss, and cognition on the pupil response. Ear and Hearing , 32(4):498–510, 2011

  205. [213]

    Weiss, Y e Jia, Zhifeng Chen, and Y onghui Wu

    Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Y e Jia, Zhifeng Chen, and Y onghui Wu. LibriTTS: A corpus derived from LibriSpeech for text-to-speech. In Proceedings of INTERSPEECH, pages 1526–1530, 2019

  206. [214]

    Instruc- tional video in e-learning: Assessing the impact of interactive video on learning effectiveness

    Dongsong Zhang, Lina Zhou, Robert O Briggs, and Jay F Nunamaker Jr. Instruc- tional video in e-learning: Assessing the impact of interactive video on learning effectiveness. Information & Management, 43(1):15–27, 2006

  207. [215]

    Video features, engagement, and patterns of collective attention allocation: An open flow network perspective

    Jingjing Zhang, Yicheng Huang, and Ming Gao. Video features, engagement, and patterns of collective attention allocation: An open flow network perspective. Journal of Learning Analytics , 9(1):32–52, 2022

  208. [216]

    WithY ou: Automated adaptive speech tutoring with context-dependent speech recognition

    Xinlei Zhang, Takashi Miyaki, and Jun Rekimoto. WithY ou: Automated adaptive speech tutoring with context-dependent speech recognition. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1–2, 2020

  209. [217]

    Embodied music training can help improve speech imitation and pronunciation skills.Language Teaching, 2024

    Yuan Zhang, Florence Baills, and Pilar Prieto. Embodied music training can help improve speech imitation and pronunciation skills.Language Teaching, 2024. Pub- lished online 18 December 2024

  210. [218]

    L2-Arctic: A non-native 105 English speech corpus

    Guanlong Zhao, Evgeny Chukharev-Hudilainen, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Ricardo Gutierrez-Osuna, and John Levis. L2-Arctic: A non-native 105 English speech corpus. In Proceedings of Interspeech 2018 , pages 2783–2787, 2018

  211. [219]

    The crowd in MOOCs: A study of learning patterns at scale

    Xin Zhou, Aixin Sun, Jie Zhang, and Donghui Lin. The crowd in MOOCs: A study of learning patterns at scale. Interactive Learning Environments , 33(3):2136– 2150, 2025

  212. [220]

    Zimmerman

    Barry J. Zimmerman. Becoming a self-regulated learner: An overview. Theory into Practice, 41(2):64–70, 2002

  213. [221]

    Effects of playback speed and language proficiency on listening comprehension of multilingual English learners

    Jiaxuan Zong, Nihat Polat, and Laura Mahalingappa. Effects of playback speed and language proficiency on listening comprehension of multilingual English learners. International Multilingual Research Journal , 2024. Published online 8 June 2024; later assigned to volume 19, iss...

  214. [222]

    DDSupport: Language Learning Support System that Displays Differences and Distances from Model Speech,

    Kazuki Kawamura and Jun Rekimoto, “DDSupport: Language Learning Support System that Displays Differences and Distances from Model Speech,” 21st IEEE International Conference on Machine Learning and Applications (ICMLA), Nas- sau, Bahamas, 2022, pp. 313-320, doi: 10.1109/ICMLA5...

  215. [223]

    AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,

    Kazuki Kawamura and Jun Rekimoto, “AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,” Augmented Hu- mans International Conference 2023 (AHs), Glasgow, United Kingdom, 2023, pp. 200-208, doi: 10.1145/3582700.3582722

  216. [224]

    FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Con- texts,

    Kazuki Kawamura and Jun Rekimoto, “FastPerson: Enhancing Video Learning through Effective Video Summarization that Preserves Linguistic and Visual Con- texts,” Augmented Humans International Conference 2024 (AHs), Melbourne, Aus- tralia, 2024, pp. 205-216, doi: 10.1145/3652920...

  217. [225]

    A Language Acquisition Support System that Presents Differences and Distances from Model Speech,

    Kazuki Kawamura and Jun Rekimoto, “A Language Acquisition Support System that Presents Differences and Distances from Model Speech,” 34th Annual ACM Symposium on User Interface Software and Technology (UIST), Virtual Event, USA, 2021, pp. 44-46, doi: 10.1145/3474349.3480225

  218. [226]

    Visualization of Speech Differences for Dialect Speech Training,

    Kazuki Kawamura and Jun Rekimoto, “Visualization of Speech Differences for Dialect Speech Training,” 2021 CHI Conference on Human Factors in Comput- ing Systems Workshop on Human Augmentation for Skill Acquisition and Skill Transfer (CHI), Virtual Event, Japan, 2021

  219. [227]

    AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,

    Kazuki Kawamura and Jun Rekimoto, “AIxSpeed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models,” 35th Annual ACM Symposium on User Interface Software and Technology (UIST), Bend, USA, 2022, pp. 1-3, doi: 10.1145/3526114.3558727

  220. [228]

    QA-FastPerson: Extending Video Platform Search Capabilities by Creating Summary Videos in Response to User Queries,

    Kazuki Kawamura and Jun Rekimoto, “QA-FastPerson: Extending Video Platform Search Capabilities by Creating Summary Videos in Response to User Queries,” Augmented Humans International Conference 2024 (AHs), Melbourne, Australia, 2024, pp. 290-293, doi: 10.1145/3652920.3653052

  221. [229]

    Generating Summary Videos from User Questions to Support Video-Based Learning,

    Kazuki Kawamura and Jun Rekimoto, “Generating Summary Videos from User Questions to Support Video-Based Learning,” 2024 CHI Conference on Human Factors in Computing Systems Workshop on Generative AI and HCI (CHI), Hon- olulu, USA, 2024. 107 Peer-Reviewed Domestic Conference Pr...

  222. [230]

    DDSupport: A Language Learning Sup- port System that Presents Differences and Distances from Model Pronunciation,

    Kazuki Kawamura and Jun Rekimoto, “DDSupport: A Language Learning Sup- port System that Presents Differences and Distances from Model Pronunciation,” Interaction, 2022, pp. 77-86 (in Japanese)

  223. [231]

    FastPerson: Lecture Video Summarization Based on Visual and Audio Information for User-Centered Learning Experience,

    Kazuki Kawamura and Jun Rekimoto, “FastPerson: Lecture Video Summarization Based on Visual and Audio Information for User-Centered Learning Experience,” Interaction, 2024, pp. 11-20 (in Japanese) 108

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.