Pith. sign in

REVIEW 3 major objections 5 minor 74 references

More AI Assistance Reduces Cognitive Engagement: Examining the AI Assistance Dilemma in AI-Supported Note-Taking

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Moderate AI assistance beats full automation for comprehension in note-taking, even though users prefer the automated tool.

desk verdict A well-run within-subject study of AI assistance levels in note-taking, but the headline contrast conflates automation level with chunk granularity and delivery cadence, so the core interpretation is not cleanly identified. read the letter →

arxiv 2509.03392 v1 pith:4IAWZRB3 submitted 2025-09-03 cs.HC

classification cs.HC
keywords AIassistancedilemmanote-takingcognitiveengagementloadhuman-AIcollaborationlearningoutcomeslargelanguagemodelswithin-subjectexperiment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether AI note-taking assistants that do more of the thinking for the user end up undermining learning. In a within-subject lab study with 30 participants watching 10-minute lectures, it compares three variants of a note-taking system: fully automated structured notes, turn-level AI summaries, and raw transcripts. The paper's central finding is that the middle level—AI summaries delivered in small blocks after each speaking turn—produced the highest post-lecture test scores, while the fully automated condition produced the lowest. Participants nevertheless rated the automated system as the most usable and said it required the least effort, revealing a gap between what people prefer and what helps them understand. The authors argue this is the "AI assistance dilemma": automating the encoding work of note-taking removes the active processing that supports comprehension, while moderate assistance preserves it.

What carries the argument

The experimental apparatus is NoteCopilot, with three assistance levels defined by how much AI pre-processes the lecture before it reaches the user: Automated AI (structured note blocks every 1–2 minutes), Intermediate AI (turn-level summary blocks, roughly every 15 seconds), and Minimal AI (transcript blocks per turn). All variants share the same interface—a real-time AI panel plus a drag-and-drop text editor—so the treatment is the granularity, timing, and abstraction of AI-generated content. The load-bearing mechanism is the degree to which AI automates the encoding stage of note-taking; the paper claims this shifts where cognitive effort goes, offloading synthesis in Automated AI (reduci

What would settle it

Run the same study with delivery cadence and block size held constant across all three conditions (e.g., every condition presents small blocks every 15 seconds, differing only in whether each block is a transcript, a summary, or a structured note). If the Intermediate advantage shrinks or disappears when cadence is equalized, the automation-level explanation is wrong; if it persists, the moderate-assistance claim is supported. Alternatively, keep the automated notes' content identical but deliver them in smaller, more frequent chunks matched to Intermediate's timing to isolate pacing from abst

Watch

Extended reading notes

Core claim

This paper establishes that in real-time AI-assisted note-taking, the amount of AI processing of lecture content has a non-monotonic effect on comprehension. Using a within-subject experiment (N=30) with three assistance levels, the authors report that students in the Intermediate AI condition (turn-level summaries arriving roughly every 15 seconds) scored significantly higher on a closed-book post-test (M = 13.88, SD = 2.34) than students in the Automated AI condition (M = 10.28, SD = 3.75), a difference of 4.17 points (p = 0.002) in a mixed-effects regression, and also outperformed Minimal AI (transcript-only blocks). After using their notes to revise answers, Intermediate AI again led, wh

Load-bearing premise

The three AI conditions differ not only in how much AI processes the content but also in how big the delivered chunks are and how often they arrive, so the comprehension gap could be caused by chunk size or timing rather than by the level of AI assistance itself.

Editorial extensions

If this is right

  • Fully automated note generation appears to reduce comprehension compared with moderate AI assistance, and using the automated notes during review does not close the gap.
  • Turn-level AI summaries that preserve user selection and integration seem to support both encoding and storage: the same condition led on closed-book and notes-allowed tests.
  • Users' preference for the most automated system, despite its lower learning outcomes, means self-reported ease and quality are unreliable proxies for cognitive benefit in AI-assisted learning tools.
  • Designers of AI assistance for learning and other cognitively demanding tasks should aim to offload mechanical effort (typing, retrieval) while preserving effortful interpretation and integration, and should not treat minimal task load as the success metric.
  • The finding suggests a concrete design target: AI assistance may be most effective when it produces small, digestible units that the user must actively assemble, rather than a final polished artifact.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The three conditions differ not only in AI abstraction level but also in block size and delivery cadence (large blocks every 1–2 min, small summary blocks every ~15 s, or per-turn transcripts), so the same data are consistent with a purely "chunking and pacing" account: smaller, more frequent blocks may drive the comprehension gain regardless of whether they are summaries or transcripts. A follow-
  • The revision-score result suggests the cost of fully automated notes is not just lost encoding; AI-structured notes may also be harder to process during review, consistent with the paper's appeal to the cost of consuming external representations. This could be tested by measuring reading time or navigation effort on AI-generated notes versus self-composed notes.
  • The paper's behavioral signal—40% modification of adopted knowledge units in the Intermediate condition—hints that edit trajectories could serve as implicit intent signals for adaptive AI, a direction the authors raise but do not implement.
  • A practical extension the paper leaves implicit: the Intermediate AI design could be deployed as an adaptive default that increases or decreases block size based on lecture density or user engagement, with a randomized study to see whether such adaptivity further improves learning outcomes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a within-subject experiment (N=30) in which participants watched short lecture videos and took notes under three conditions: Automated AI (structured AI note blocks every 1–2 minutes), Intermediate AI (turn-level summary blocks approximately every 15 seconds), and Minimal AI (turn-level transcript blocks). The authors measure post-test comprehension (closed-book and note-revision scores), cognitive load, note-taking behaviors, usability, and qualitative perceptions. The central quantitative result is that Intermediate AI significantly outperforms Automated AI on both post-test scores (coefficient 4.17, p=0.002) and revision scores (coefficient 2.85, p=0.017), while participants nevertheless prefer Automated AI for ease of use and perceived quality. The paper interprets these findings as evidence for the 'AI assistance dilemma' in note-taking: over-automation reduces cognitive engagement and comprehension, whereas moderate assistance preserves beneficial encoding.

Significance. If the central claim were cleanly identified, this would be a valuable contribution to CSCW and HCI. The study has notable strengths: a within-subject design with counterbalanced condition order, a video covariate in the mixed-effects models, pre-generated AI content to keep stimulus content identical across participants, blinded grading with reported intercoder reliability, and behavioral logging that goes beyond self-report. The finding that users prefer the condition with the worst learning outcome is practically important and aligns with prior work on effort minimization and over-reliance on AI. The qualitative data also provide rich, plausible accounts of how intermediate summaries support active selection and integration. However, the main causal interpretation is currently threatened by a construct-validity confound, and the abstract overstates one pairwise result. These issues are load-bearing for the paper's headline claim, so the manuscript needs revision rather than acceptance in its current form.

major comments (3)
  1. [§3.2.2, §3.2.3, Table 2] The Intermediate-vs-Automated contrast, which is the paper's main quantitative result (Table 2, coefficient 4.172, p=0.002), manipulates more than the level of AI processing. Per §3.2.2, Automated AI emits one structured block every 1–2 minutes, Intermediate AI emits a summary every ~15 seconds, and Minimal AI emits transcript blocks per turn. The control claim in §3.2.3 only equalizes transcript and summary segmentation by natural turns; Automated AI deliberately combines several turns into topic-based blocks. Thus block length and delivery cadence change together with abstraction level. The Intermediate advantage could be due to smaller, more frequent chunks or lower working-memory cost per integration event rather than to 'moderate AI assistance' as a construct. Section 5.2.1 even quotes participants attributing the benefit to 'digestible chunks' and a 'constant flow,' which are granu
  2. [Abstract, §5.1.1, Table 2] The abstract and Figure 1 state that Automated AI 'yielded the lowest' post-test scores. This is not consistently supported by the reported statistics. For post-test scores, the mixed model shows Minimal AI significantly above Automated AI (coefficient 2.702, p=0.044), but the post-hoc Tukey HSD does not report a significant Minimal-vs-Automated difference; it only reports Intermediate-vs-Automated and Intermediate-vs-Minimal. For revision scores, Minimal AI is not significantly different from Automated AI (coefficient 0.367, p=0.763). Thus Automated AI is significantly worse than Intermediate AI, but calling it 'the lowest' overstates a pairwise result that is inconsistent across models and outcome measures. The claims in §6 and §7 should be qualified accordingly.
  3. [§3.1, §5.1.1] The paper's causal language ('More AI Assistance Reduces Cognitive Engagement') implies a monotonic effect of assistance level, but the study lacks a no-AI baseline and the Minimal AI condition is itself an AI-assistance condition. More importantly, the paper does not directly measure cognitive engagement; it infers engagement from comprehension scores and from some behavioral proxies. This is reasonable as a signal, but the title and RQ framing should be careful to distinguish 'AI assistance level as operationalized here' from 'AI assistance in general.' The absence of a no-AI control is acknowledged in §3.2.2, but the framing in the title and abstract should match the actual comparisons.
minor comments (5)
  1. [§4.1.2, Table 3] The extraneous-load items are labeled with duplicate numbers: two items are labeled Q5 and the subsequent reading-effort item is labeled Q6. Please renumber the items consistently.
  2. [§5.1.3, Table 5] There is an inconsistency in the reported mean for Automated AI in the normalized AI-content drops analysis: the text states Mean = 9.41 (SD = 4.89), while Table 5 reports Mean = 10.03 (SD = 4.89). Please correct and ensure the reported p-value for the Intermediate-vs-Automated comparison corresponds to the correct descriptive statistics.
  3. [§4.1.3] Typos: 'qantity' and 'qality' should be 'quantity' and 'quality.' Also, the limitation about not measuring note quality is acknowledged but appears abruptly; consider moving it to the limitations paragraph in §6.4.
  4. [§3.1.3] The paper states that videos 'required no prior knowledge' and that no pre-test was included because videos were matched on difficulty. Since the mixed-effects model showed no significant video effect in §5.1.1, this choice is defensible, but a sentence reporting the pilot validation of comparable difficulty would strengthen the justification.
  5. [§5.2.5] The sentence 'Interestingly, some students used the search feature not just to fill gaps but also to verify the accuracy of AI-generated notes' is followed by a quote from P5. The quote illustrates a trust verification episode, but the connection to the preceding claim could be made more explicit.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; the central result is an empirical contrast from new data. Minor self-citations are not load-bearing.

full rationale

This paper is an empirical within-subject experiment, not a formal derivation or prediction exercise. The claimed reasoning chain is: prior theory (encoding-storage, assistance dilemma) motivates hypotheses; three conditions are operationalized; post-test and revision scores are measured; mixed-effects regression summarizes the observed contrasts. No step in this chain defines the outcome in terms of the manipulation by construction. The key quantitative result in Section 5.1.1 (Intermediate AI coefficient 4.17, p = 0.002) is estimated from the observed post-test scores in Table 2; it is a data summary, not a fitted parameter that is then renamed as a prediction. The design draws on prior work, including some self-citations (e.g., MeetMap [10], Meetscript [8], and related papers), but the central claim does not reduce to accepting those citations: the outcome is independently measured on 30 participants. The more serious concern is construct validity, not circularity: Section 3.2.3 claims granularity and completeness are controlled, but Automated AI (1-2 minute structured blocks), Intermediate AI (~15 second summary blocks), and Minimal AI (turn-level transcript blocks) vary in abstraction, block size, and delivery cadence simultaneously. That confound threatens the causal interpretation of the automation level, but it does not make the result equivalent to its inputs by definition. No passage in the manuscript asserts that a derivation is circular or omits a proof that would be needed for a self-definitional argument. A score of 1 reflects the presence of minor, non-load-bearing self-citations in the related work and design rationale, with no circular dependence in the empirical chain.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the construct validity of the three assistance levels, the equivalence of the three videos, and the validity of the post-test. None of these is externally benchmarked; they are design assumptions.

free parameters (3)
  • Automated AI note-block interval = 90 seconds
    Hand-chosen in §3.2.2 to mirror commercial AI note tools; the effect size may depend on this cadence.
  • Intermediate AI summary-block interval = ~15 seconds
    Chosen to match natural sentence turns; a different cadence could change the result.
  • Post-test item weights = 2 MCQs + 3 OEQs, 4 points each OEQ
    Rubric designed by authors, piloted with 4 students; no external validation of the measure's sensitivity.
assumptions (4)
  • domain assumption Encoding-storage paradigm (Kiewra, 1989) explains note-taking effects
    Used throughout the Discussion to interpret results; not independently tested in this study.
  • domain assumption Post-test scores measure comprehension
    The post-test was expert-authored and piloted, but no reliability/validity evidence (e.g., test-retest, correlation with external measures) is provided.
  • domain assumption The three lecture videos are comparable in difficulty and require no prior knowledge
    Stated in §3.1.3; the authors decided against a pre-test because they assumed comparability. The regression found no significant video effect, but with N=30 the test is low-powered.
  • domain assumption The AI-generated notes in all three conditions are equally accurate
    Intermediate and Automated notes were manually reviewed by researchers, but no quantitative accuracy check is reported; the Minimal condition uses raw transcripts, so accuracy differences could confound the comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of More AI Assistance Reduces Cognitive Engagement: Examining the AI Assistance Dilemma in AI-Supported Note-Taking." pith.science (2026). https://pith.science/paper/4IAWZRB3

@misc{pith2026250903392,
  author       = {Pith},
  title        = {Pith review of: More AI Assistance Reduces Cognitive Engagement: Examining the AI Assistance Dilemma in AI-Supported Note-Taking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4IAWZRB3}},
  note         = {Machine review of arXiv:2509.03392}
}
read the original abstract

As AI tools become increasingly embedded in cognitively demanding tasks such as note-taking, questions remain about whether they enhance or undermine cognitive engagement. This paper examines the "AI Assistance Dilemma" in note-taking, investigating how varying levels of AI support affect user engagement and comprehension. In a within-subject experiment, we asked participants (N=30) to take notes during lecture videos under three conditions: Automated AI (high assistance with structured notes), Intermediate AI (moderate assistance with real-time summary, and Minimal AI (low assistance with transcript). Results reveal that Intermediate AI yields the highest post-test scores and Automated AI the lowest. Participants, however, preferred the automated setup due to its perceived ease of use and lower cognitive effort, suggesting a discrepancy between preferred convenience and cognitive benefits. Our study provides insights into designing AI assistance that preserves cognitive engagement, offering implications for designing moderate AI support in cognitive tasks.

Figures

Figures reproduced from arXiv: 2509.03392 by the authors.

Figure 1
Figure 1. In a within-subject experiment with 30 participants, we tested three variants of an AI-assisted note [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Study Design. The study started with a 5-minute introduction outlining its objectives. Participants then experienced all three AI assistance conditions— Automated AI , Intermediate AI , Minimal AI , in a specified order. Each session began with a demo instruction and a toy task. Participants watched a 10-minute video while taking notes with the system, after which they had 1 minute to organize their notes. This was … view at source ↗
Figure 3
Figure 3. System Overview. On the upper left is the Real-time AI-Generated Note panel (a), where users can see AI-generated notes and click on them to see related transcripts (a1). On the lower left is the Search and Synthesis Panel (b), where users can interact with AI by requesting notes centered around a topic (b1) based on previous course content.(b2). On the right is the rich Text-Editor (c), where users can create their… view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Automated AI and Minimal AI Interface. On the left is the Automated AI Interface (n), where users can drag and drop AI-generated notes alongside related transcripts from the Real-time AI-Generated Note panel (n1) and Search and Synthesis Panel (n2) into the Text-Editor…
Figure 5
Figure 5. Figure 5: The granularity of AI notes in the three conditions. From left to right, the AI-generated results of three systems— Minimal AI , Intermediate AI , and Automated AI —illustrate the topic of forced display with increasing levels of AI assistance. The Minimal AI system pr…
Figure 6
Figure 6. Figure 6: Post-hoc Tukey HSD analysis comparing post-test and revision scores across conditions. The left bar graph shows the Post-test Scores, where the Intermediate AI condition significantly outperformed both the Automated AI and Minimal AI . The right bar graph displays the …
Figure 7
Figure 7. Figure 7: Comparison of note-taking behaviors across conditions with varying levels of AI assistance. The left chart shows the average total note count for each condition, divided into AI-generated content and manually typed content, with error bars indicating standard deviation…
Figure 8
Figure 8. Figure 8: Usability Rating. This figure presents the mean usability ratings across three conditions— [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 60 canonical work pages

  1. [1]

    Sumit Asthana, Sagih Hilleli, Pengcheng He, and Aaron Halfaker. 2023. Summaries, Highlights, and Action items: Design, implementation and evaluation of an LLM-powered meeting recap system. arXiv preprint arXiv:2307.15793 (2023)

  2. [2]

    Roger Azevedo, Jennifer G Cromley, and Diane Seibert. 2004. Does adaptive scaffolding facilitate students’ ability to regulate their learning with hypermedia? Contemporary educational psychology 29, 3 (2004), 344–370

  3. [3]

    Aaron Bauer and Kenneth R Koedinger. 2007. Selection-based note-taking applications. In Proceedings of the SIGCHI conference on Human factors in computing systems . 981–990

  4. [4]

    Elizabeth L Bjork, Robert A Bjork, et al. 2011. Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. Psychology and the real world: Essays illustrating fundamental contributions to society 2, 59-68 (2011)

  5. [5]

    Virginia Braun and Victoria Clarke. 2012. Thematic analysis. American Psychological Association

  6. [6]

    Bui, Joel Myerson, and Sandra Hale

    Dung C. Bui, Joel Myerson, and Sandra Hale. 2013. Note-taking with computers: Exploring alternative strategies for improved recall. Journal of Educational Psychology 105, 2 (2013), 299–309. https://doi.org/10.1037/a0030367

  7. [7]

    Ting-Ju Chen. 2021. Association, Reflection, Stimulation: Problem Exploration in Early Design through AI-Augmented Mind-Mapping. Thesis. https://oaktrust.library.tamu.edu/handle/1969.1/195385 Accepted: 2022-01-27T22:18:02Z

  8. [8]

    Xinyue Chen, Shuo Li, Shipeng Liu, Robin Fowler, and Xu Wang. 2023. Meetscript: designing transcript-based interactions to support active participation in group video meetings. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2 (2023), 1–32. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW451. Publication date: November 2025. AI As...

Show all 74 references
  1. [9]

    Xinyue Chen, Lev Tankelevitch, Rishi Vanukuru, Ava Elizabeth Scott, Payod Panda, and Sean Rintel. 2025. Are We On Track? AI-Assisted Active and Passive Goal Reflection During Meetings. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–22

  2. [10]

    Xinyue Chen, Nathan Yap, Xinyi Lu, Aylin Gunal, and Xu Wang. 2025. MeetMap: Real-Time Collaborative Dialogue Mapping with LLMs in Online Meetings. Proceedings of the ACM on Human-Computer Interaction 9, 2 (2025), 1–35

  3. [11]

    Jamie Costley, Matthew Courtney, and Mik Fanguy. 2022. The interaction of collaboration, note-taking completeness, and performance over 10 weeks of an online course. The Internet and Higher Education 52 (2022), 100831

  4. [12]

    Jamie Costley and Mik Fanguy. 2021. Collaborative note-taking affects cognitive load: the interplay of completeness and interaction. Educational Technology Research and Development 69 (2021), 655–671

  5. [13]

    Richard C Davis, James A Landay, Victor Chen, Jonathan Huang, Rebecca B Lee, Frances C Li, James Lin, Charles B Morrey III, Ben Schleimer, Morgan N Price, et al. 1999. NotePals: Lightweight note sharing by the group, for the group. In Proceedings of the SIGCHI conference on Hu...

  6. [14]

    Paramveer S Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel Peter Robert. 2024. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. In Proceedings of the CHI Conference on Human Factors in Computing Syst...

  7. [15]

    Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel P

    Paramveer S. Dhillon, Somayeh Molaei, Jiaqi Li, Maximilian Golub, Shaochun Zheng, and Lionel P. Robert. 2024. Shaping Human-AI Collaboration: Varied Scaffolding Levels in Co-writing with Language Models. http://arxiv.org/ abs/2402.11723 arXiv:2402.11723 [cs]

  8. [16]

    Jingchao Fang, Yanhao Wang, Chi-Lan Yang, Ching Liu, and Hao-Chuan Wang. 2022. Understanding the Effects of Structured Note-taking Systems for Video-based Learners in Individual and Social Learning Contexts. Proceedings of the ACM on Human-Computer Interaction 6, GROUP (Jan. 2...

  9. [17]

    Jingchao Fang, Yanhao Wang, Chi-Lan Yang, and Hao-Chuan Wang. 2021. NoteCoStruct: Powering online learners with socially scaffolded note taking and sharing. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. 1–5

  10. [18]

    Mik Fanguy, Matthew Baldwin, Evgeniia Shmeleva, Kyungmee Lee, and Jamie Costley. 2023. How collaboration influences the effect of note-taking on writing performance and recall of contents. Interactive Learning Environments 31, 7 (Oct. 2023), 4057–4071. https://doi.org/10.1080/...

  11. [19]

    Flanigan and Scott Titsworth

    Abraham E. Flanigan and Scott Titsworth. 2020. The impact of digital distraction on lecture note taking and student learning. Instructional Science 48, 5 (Oct. 2020), 495–524. https://doi.org/10.1007/s11251-020-09517-2

  12. [20]

    Krzysztof Z Gajos and Lena Mamykina. 2022. Do people engage cognitively with AI? Impact of AI assistance on incidental learning. In Proceedings of the 27th International Conference on Intelligent User Interfaces . 794–806

  13. [21]

    Jie Gao, Kenny Tsu Wei Choo, Junming Cao, Roy Ka-Wei Lee, and Simon Perrault. 2023. CoAIcoder: Examining the effectiveness of AI-assisted human-to-human collaboration in qualitative analysis. ACM Transactions on Computer- Human Interaction 31, 1 (2023), 1–38

  14. [22]

    Nitesh Goyal, Gilly Leshed, and Susan R. Fussell. 2013. Effects of visualization and note-taking on sensemaking and analysis. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . ACM, Paris France, 2721–2724. https://doi.org/10.1145/2470654.2481376

  15. [23]

    Ziwei Gu, Ian Arawjo, Kenneth Li, Jonathan K Kummerfeld, and Elena L Glassman. 2024. An AI-Resilient Text Rendering Technique for Reading and Skimming Documents. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–22

  16. [24]

    Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904–908

  17. [25]

    Ken Hinckley, Shengdong Zhao, Raman Sarin, Patrick Baudisch, Edward Cutrell, Michael Shilman, and Desney Tan

  18. [26]

    Xinying Hou, Barbara Jane Ericson, and Xu Wang. 2022. Using Adaptive Parsons Problems to Scaffold Write-Code Problems. In Proceedings of the 2022 ACM Conference on International Computing Education Research - Volume 1 . ACM, Lugano and Virtual Event Switzerland, 15–26. https:/...

  19. [27]

    Xinying Hou, Zihan Wu, Xu Wang, and Barbara J Ericson. 2025. Personalized Parsons Puzzles as Scaffolding Enhance Practice Engagement Over Just Showing LLM-Powered Solutions. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 2 . 1483–1484

  20. [28]

    Faria Huq, Abdus Samee, David Chuan-en Lin, Xiaodi Alice Tang, and Jeffrey P Bigham. 2024. NoTeeline: Supporting Real-Time Notetaking from Keypoints with Large Language Models. arXiv preprint arXiv:2409.16493 (2024)

  21. [29]

    Renée S Jansen, Daniel Lakens, and Wijnand A IJsselsteijn. 2017. An integrative review of the cognitive costs and benefits of note-taking. Educational Research Review 22 (2017), 223–233

  22. [30]

    Slava Kalyuga. 2009. Adapting levels of instructional support to optimize learning complex cognitive skills. In Managing Cognitive Load in Adaptive Multimedia Learning . IGI Global, 246–271. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW451. Publication date: No...

  23. [31]

    Matthew Kam, Jingtao Wang, Alastair Iles, Eric Tse, Jane Chiu, Daniel Glaser, Orna Tarshish, and John Canny. 2005. Livenotes: a system for cooperative and augmented note-taking in lectures. In Proceedings of the SIGCHI conference on Human factors in computing systems . 531–540

  24. [32]

    Kelly, Jason M

    Anam Ahmad Khan, Sadia Nawaz, Joshua Newn, Ryan M. Kelly, Jason M. Lodge, James Bailey, and Eduardo Velloso

  25. [33]

    Kenneth A. Kiewra. 1989. A review of note-taking: The encoding-storage paradigm and beyond.Educational Psychology Review 1, 2 (June 1989), 147–172. https://doi.org/10.1007/BF01326640

  26. [34]

    Kiewra, Nelson F

    Kenneth A. Kiewra, Nelson F. DuBois, David Christian, Anne McShane, Michelle Meyerhoffer, and David Roskelley

  27. [35]

    Melina Klepsch, Florian Schmitz, and Tina Seufert. 2017. Development and validation of two instruments measuring intrinsic, extraneous, and germane cognitive load. Frontiers in psychology 8 (2017), 294028

  28. [36]

    Koedinger and Vincent Aleven

    Kenneth R. Koedinger and Vincent Aleven. 2007. Exploring the Assistance Dilemma in Experiments with Cognitive Tutors. Educational Psychology Review 19, 3 (Sept. 2007), 239–264. https://doi.org/10.1007/s10648-007-9049-0

  29. [37]

    Kenneth R Koedinger, Albert T Corbett, and Charles Perfetti. 2012. The Knowledge-Learning-Instruction framework: Bridging the science-practice chasm to enhance robust student learning. Cognitive science 36, 5 (2012), 757–798

  30. [38]

    Anastasia Kuzminykh and Sean Rintel. 2020. Classification of Functional Attention in Video Meetings. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . ACM, Honolulu HI USA, 1–13. https://doi.org/10. 1145/3313831.3376546

  31. [39]

    Hui Yi Leong, Yi Fan Gao, Shuai Ji, Bora Kalaycioglu, and Uktu Pamuksuz. 2024. A GEN AI Framework for Medical Note Generation. arXiv preprint arXiv:2410.01841 (2024)

  32. [40]

    Jimmie Leppink, Fred Paas, Cees PM Van der Vleuten, Tamara Van Gog, and Jeroen JG Van Merriënboer. 2013. Development of an instrument for measuring different types of cognitive load. Behavior research methods 45 (2013), 1058–1072

  33. [41]

    Susan Lin, Jeremy Warner, JD Zamfirescu-Pereira, Matthew G Lee, Sauhard Jain, Shanqing Cai, Piyawat Lertvit- tayakumjorn, Michael Xuelin Huang, Shumin Zhai, Björn Hartmann, et al. 2024. Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation. In Proceedings ...

  34. [42]

    Ching Liu, Chi-Lan Yang, Joseph Jay Williams, and Hao-Chuan Wang. 2019. Notestruct: Scaffolding note-taking while learning from online videos. In Extended abstracts of the 2019 CHI conference on human factors in computing systems . 1–6

  35. [43]

    Tamas Makany, Jonathan Kemp, and Itiel E. Dror. 2009. Optimising the use of note-taking as an external cognitive aid for increasing learning. British Journal of Educational Technology 40, 4 (July 2009), 619–635. https://doi.org/10.1111/j.1467- 8535.2008.00906.x

  36. [44]

    Bruce M McLaren, S Lim, and Kenneth R Koedinger. 2008. When and how often should worked examples be given to students? New results and a summary of the current state of research. In Proceedings of the 30th annual conference of the cognitive science society . 2176–2181

  37. [45]

    Andrew M Mcnutt, Chenglong Wang, Robert A Deline, and Steven M. Drucker. 2023. On the Design of AI-powered Code Assistants for Notebooks. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM, Hamburg Germany, 1–16. https://doi.org/10.1145/35445...

  38. [46]

    Fred GWC Paas and Jeroen JG Van Merriënboer. 1993. The efficiency of instructional conditions: An approach to combine mental effort and performance measures. Human factors 35, 4 (1993), 737–743

  39. [47]

    Annie Piolat, Thierry Olive, and Ronald Kellogg. 2005. Cognitive effort during note taking.Applied Cognitive Psychology 19 (April 2005), 291–312. https://doi.org/10.1002/acp.1086 270 citations (Crossref) [2024-01-03]

  40. [48]

    Yi Ren, Yang Li, and Edward Lank. 2014. InkAnchor: enhancing informal ink-based note taking on touchscreen mobile phones. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 1123–1132

  41. [49]

    Ido Roll, Vincent Aleven, Bruce M McLaren, and Kenneth R Koedinger. 2007. Designing for metacognition—applying cognitive tutor principles to the tutoring of help seeking. Metacognition and Learning 2 (2007), 125–140

  42. [50]

    Ido Roll, Vincent Aleven, Bruce M McLaren, and Kenneth R Koedinger. 2011. Improving students’ help-seeking skills using metacognitive feedback in an intelligent tutoring system. Learning and instruction 21, 2 (2011), 267–280

  43. [51]

    Rong, Mo Morgana Zhou, Ge Gao, and Zhicong Lu

    Ethan Z. Rong, Mo Morgana Zhou, Ge Gao, and Zhicong Lu. 2023. Understanding Personal Data Tracking and Sensemaking Practices for Self-Directed Learning in Non-classroom and Non-computer-based Contexts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Sys...

  44. [52]

    Russell, Mark J

    Daniel M. Russell, Mark J. Stefik, Peter Pirolli, and Stuart K. Card. 1993. The cost structure of sensemaking. In Proceedings of the SIGCHI conference on Human factors in computing systems - CHI ’93 . ACM Press, Amsterdam, The Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, A...

  45. [53]

    Martin Schroder. 2023. AutoScrum: Automating Project Planning Using Large Language Models. https://doi.org/10. 48550/arXiv.2306.03197 arXiv:2306.03197 [cs]

  46. [54]

    Ava Elizabeth Scott, Lev Tankelevitch, Payod Panda, Rishi Vanukuru, Xinyue Chen, and Sean Rintel. 2025. What Does Success Look Like? Catalyzing Meeting Intentionality with AI-Assisted Prospective Reflection. In Proceedings of the 4th Annual Symposium on Human-Computer Interact...

  47. [55]

    Ava Elizabeth Scott, Lev Tankelevitch, and Sean Rintel. 2024. Mental Models of Meeting Goals: Supporting Intentionality in Meeting Technologies. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  48. [56]

    Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. 2024. AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–17. https://doi...

  49. [57]

    Meng Sun, Minhong Wang, Rupert Wegerif, and Jun Peng. 2022. How do students generate ideas together in scientific creativity tasks through computer-based mind mapping? Computers & Education 176 (Jan. 2022), 104359. https://doi.org/10.1016/j.compedu.2021.104359 20 citations (Cr...

  50. [58]

    Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel

  51. [59]

    Johanna Telenius. 2016. Sensemaking in Meetings - Collaborative Construction of Meaning and Decisions through Epistemic Authority. Aalto University. https://aaltodoc.aalto.fi/handle/123456789/23391 ISSN: 1799-4942 (electronic)

  52. [60]

    Hsin-Ruey Tsai, Shih-Kang Chiu, and Bryan Wang. 2024. GazeNoter: Co-Piloted AR Note-Taking via Gaze Selection of LLM Suggestions to Match Users’ Intentions. arXiv preprint arXiv:2407.01161 (2024)

  53. [61]

    Karthikeyan Umapathy. [n. d.]. Requirements to support Collaborative Sensemaking. ([n. d.])

  54. [62]

    Jixuan Wang, Jingbo Yang, Haochi Zhang, Helen Lu, Marta Skreta, Mia Husić, Aryan Arbabi, Nicole Sultanum, and Michael Brudno. 2022. PhenoPad: building AI enabled note-taking interfaces for patient encounters. NPJ digital medicine 5, 1 (2022), 12

  55. [63]

    Ruotong Wang, Lin Qiu, Justin Cranshaw, and Amy X. Zhang. 2024. Meeting Bridges: Designing Information Artifacts that Bridge from Synchronous Meetings to Asynchronous Collaboration. http://arxiv.org/abs/2402.03259 arXiv:2402.03259 [cs]

  56. [64]

    Witherby and Sarah K

    Amber E. Witherby and Sarah K. Tauber. 2019. The Current Status of Students’ Note-Taking: Why and How Do Students Take Notes? Journal of Applied Research in Memory and Cognition 8, 2 (June 2019), 139–153. https: //doi.org/10.1016/j.jarmac.2019.04.002

  57. [65]

    Sarah Shi Hui Wong and Stephen Wee Hun Lim. 2023. Take notes, not photos: Mind-wandering mediates the impact of note-taking strategies on video-recorded lecture learning performance. Journal of Experimental Psychology: Applied 29, 1 (March 2023), 124–135. https://doi.org/10.10...

  58. [66]

    David Wood. 2001. Scaffolding, contingent tutoring, and computer-supported learning. International Journal of Artificial Intelligence in Education 12, 3 (2001), 280–293

  59. [67]

    Xiaotong Xu, Jiayu Yin, Catherine Gu, Jenny Mar, Sydney Zhang, Jane L E, and Steven P Dow. 2024. Jamplate: exploring llm-enhanced templates for idea reflection. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 907–921

  60. [68]

    Zheng Zhang, Jie Gao, Ranjodh Singh Dhaliwal, and Toby Jia-Jun Li. 2023. Visar: A human-ai argumentative writing assistant with visual programming and rapid draft prototyping. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–30

  61. [69]

    Zheng Zhang, Weirui Peng, Xinyue Chen, Luke Cao, and Toby Jia-Jun Li. 2025. LADICA: a large shared display interface for generative AI cognitive assistance in co-located team collaboration. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–22

  62. [70]

    Rebecca Zheng, Marina Fernández Camporro, Hugo Romat, Nathalie Henry Riche, Benjamin Bach, Fanny Chevalier, Ken Hinckley, and Nicolai Marquardt. 2021. Sketchnote components, design space dimensions, and strategies for effective visual note taking. In Proceedings of the 2021 CH...

  63. [1991]

    Journal of Educational Psychology 83, 2 (June 1991), 240–245

    Note-taking functions and techniques. Journal of Educational Psychology 83, 2 (June 1991), 240–245. https: //doi.org/10.1037/0022-0663.83.2.240 Publisher: American Psychological Association

  64. [2007]

    In Proceedings of the SIGCHI conference on human factors in computing systems

    InkSeine: In Situ search for active note taking. In Proceedings of the SIGCHI conference on human factors in computing systems. 251–260

  65. [2022]

    In CHI Conference on Human Factors in Computing Systems

    To type or to speak? The effect of input modality on text understanding during note-taking. In CHI Conference on Human Factors in Computing Systems . ACM, New Orleans LA USA, 1–15. https://doi.org/10.1145/3491102.3501974

  66. [2024]

    In Proceedings of the CHI Conference on Human Factors in Computing Systems

    The metacognitive demands and opportunities of generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–24

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.