{"id":"06f13fef-2aa5-4c24-b40d-8e9a92b29946","arxiv_id":"2411.09879","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MEMA is a publicly released multi-label EEG dataset for online-learning attention classification, with baseline accuracies up to 85.12% subject-dependent and 64.84% cross-subject.","lead":"This paper presents MEMA, a new public EEG dataset for classifying attention states in online learning, with recordings from 20 people across three attention conditions and 1,060 minutes of data. It adds emotion, personality, and personal information labels so attention can be studied alongside other psychological states.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Attention labels in MEMA are conflated with the specific video stimuli and unaudited self-reports; without evidence of cross-content generalization, the Table II accuracies cannot establish a validated attention-state dataset.","rationale":"The reader's CONDITIONAL verdict is appropriate. My stress-test narrowed the concern to a single testable assumption: the three task videos are the only operationalization of attention, and no independent evidence shows that the EEG signal encodes attention rather than video identity. This is not a disagreement with the community's general practice of using task-induced states; it is a correctness risk specific to a dataset whose claimed contribution is validated labels. The internal checks in the paper (one-subject alpha/beta plots, Chi-square emotion correlations, multi-task results) are supportive but partly circular, because the emotion self-reports are collected in the same session immediately after the same videos. I do not see a need to change the reader's verdict: the paper should remain CONDITIONAL on providing label-validation evidence (self-report/task agreement, quiz performance, and/or multi-video generalization). No formal verification or parameter-free derivation is claimed, so the empirical validation burden is exactly where this concern sits.","tokens_in":7933,"tokens_out":5930,"duration_ms":75325,"concrete_test":"Recruit a subset of the original 20 participants (at least 10) and record two additional video instances per attention condition that are matched in duration and task instructions but differ in content (e.g., two different relaxing nature videos, two different lecture topics, two different blank-screen variants). Train the same EEGNet pipeline exactly as in Table II on the original MEMA trials, then evaluate it on the new sessions under the cross-subject protocol. If the mean three-class accuracy is not significantly above the 33.3% chance level, the original accuracies are explained by video-specific stimulus confounds rather than by attention-state information. If accuracy remains near the original 64.84%, the central claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that MEMA is a validated EEG dataset whose attention labels correspond to mental attention states rather than to the particular materials shown. In Section II.C, each attention state is realized by exactly one video: neutral = one-minute blank screen, relaxing = a five-minute nature-scenery video with music, concentrating = a five-minute machine-learning lecture clip. The attention label is assigned from the task and then only corroborated by a 15-second self-assessment after each video; no objective behavioral measure (e.g., quiz accuracy, reaction time, eye tracking) is reported, and the quiz answers mentioned in the concentrating task are not analyzed. This creates two correlated threats. First, the label may simply reproduce the task condition if subjects infer the intended state and answer accordingly, so self-reports are not independent validation. Second, and more fundamentally, the setup confounds attention state with low-level stimulus identity (static blank screen vs. scenic visuals and music vs. talking-head lecture). The high classification accuracies in Table II — up to 85.12% subject-dependent and 64.84% cross-subject — could therefore reflect classifiers recognizing which video is being watched through stimulus-locked EEG rather than recognizing a generalizable internal attention state. The one-subject alpha/beta power plots in Section III.B are descriptive and do not validate the labels. Because the dataset's scientific value rests on the labels, this unverified equivalence is the weakest load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MEMA, a public multi-label EEG dataset for mental attention state classification in online learning. Twenty subjects completed 12 trials each across three attention states (neutral, relaxing, concentrating), with auxiliary emotional VAD labels, Big Five personality data, and personal information. The authors provide baselines for subject-dependent and cross-subject classification of attention and emotion using classical and deep learning models, reporting attention accuracies up to 85.12% and 64.84%, respectively, and a multi-label correlation analysis between attention and emotion. The central claim is that MEMA is a validated, high-quality resource that fills a gap in public EEG attention datasets.","tokens_in":8171,"tokens_out":5286,"duration_ms":49326,"significance":"If the attention labels indeed reflect the intended mental states, MEMA would be a valuable public resource, being one of the few public EEG datasets for attention in online learning and the only one with multi-label emotion annotations, personality traits, and personal information. The paper ships a publicly available dataset and a reproducible baseline suite, including both classical and deep learning models, which is a practical strength. However, the validation claim rests on the assumptions that the task stimuli evoke separable attention states and that self-reports are reliable; the current evidence does not conclusively establish these, so the dataset's scientific value is present but not yet fully demonstrated.","major_comments":[{"comment":"The three attention states are each instantiated by exactly one video type: a one-minute blank screen for neutral, a five-minute nature-scenery video with music for relaxing, and a five-minute machine-learning lecture for concentrating. Because the stimuli differ in low-level visual content, audio content, and duration, the high classification accuracies in Table II could reflect classifiers recognizing the specific stimulus rather than a generalizable internal attention state. The paper's central claim that MEMA is a validated attention-state dataset requires evidence of cross-content generalization, such as multiple videos per state or an analysis that removes stimulus-locked features; as it stands, this confounding is unresolved.","section":"Section II.C"},{"comment":"The attention label is determined by the task condition and only corroborated by a 15-second self-assessment after each video; no objective behavioral measure (quiz accuracy, reaction time, eye tracking) is reported, and the quiz answers from the concentrating task are not analyzed. Because participants were told which state each video was intended to induce, their self-reports may reflect demand characteristics rather than independent verification. This is load-bearing for the dataset's label validity, and the authors should either provide an objective validation of the self-reports or explicitly document and discuss the limitation.","section":"Section II.C"},{"comment":"Table II reports only mean accuracy and F1 without standard deviations, confidence intervals, or significance tests. The subject-dependent setup uses one fixed split (first 9 trials for training, last 3 for testing), which is highly sensitive to trial order and within-session fatigue or learning effects. For a dataset paper, these numbers are the principal quantitative evidence of data quality; the absence of variability measures and statistical comparisons makes the evidence incomplete. Standard deviations across subjects/folds, per-class results, and chance-level comparisons should be reported.","section":"Section III.C.1 and Table II"},{"comment":"The qualitative validation in Section III.B is performed on one subject only (Figure 2). The observation that alpha power increases and beta power decreases with lower attention, while consistent with the literature, is not established across the 20-subject cohort and therefore does not, by itself, support the dataset-level quality claim. The authors should either extend the qualitative analysis to multiple subjects or clearly present it as an illustrative example rather than part of the validation.","section":"Section III.B"}],"minor_comments":[{"comment":"The dataset URL in the abstract (https://github.com/GuanjianLiu/MEMA) differs from the URL in the full text (https://github.com/XJTU-EEG/MEMA); the authors should ensure a single working link is used consistently.","section":"Abstract vs. full text"},{"comment":"The heading 'Task Design for Each Trail' should read 'Each Trial'.","section":"Section II.C"},{"comment":"The sentence 'the first 9 of one subject's trials used for training' is missing a verb; it should be 'were used for training'.","section":"Section III.C.1"},{"comment":"Table III reports chi-square values without degrees of freedom or p-values; these should be added to support the stated association claims.","section":"Table III"},{"comment":"Reference [21] is cited for the duration of the concentrating task through the Continuous Performance Test, but [21] concerns visual sustained attention degradation, not the CPT; please verify the citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main concern is whether the dataset validates attention states rather than stimulus identity. If the authors cannot add new data, they should at least reposition the paper's claims and provide explicit limitations. The reviewer sees the dataset as potentially useful even with these caveats, so rejection is not warranted, but the validation language in the abstract and conclusion is currently too strong."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: MEMA is a genuinely useful public resource—20 subjects, 1,060 minutes of EEG with attention, emotion, personality, and demographic labels, all in an online-learning context. That combination is scarce. The paradigm is clearly described, the baseline battery is standard, and the authors make the data publicly available. If you need an EEG dataset with multiple label types from a learning setting, this is worth a look.\n\nThe soft spots are real and load-bearing. Each attention state is paired with exactly one video type: neutral = blank screen, relaxing = nature scenery with music, concentrating = a machine-learning lecture. So the 'state' is perfectly confounded with the stimulus. The high subject-dependent accuracies (up to 85%+) might largely reflect classifiers recognizing which video is playing from stimulus-locked EEG, not a generalizable internal attention state. The self-reported attention labels are the only check, and the quiz answers from the concentrating task are never analyzed. That leaves the central claim—that this is a validated dataset of attention states—under-supported.\n\nThe missing standard deviations and significance tests in Table II are a smaller issue, as is the fixed train/test split in the subject-dependent setup. The qualitative validation covers one subject, which is suggestive but not strong. There's also a trivial URL mismatch between the abstract and the full text.\n\nThat said, none of this is fatal. The dataset can still be useful if the authors 1) acknowledge the stimulus confound and either add a second neutral/relaxing/concentrating video per state or explicitly frame the labels as task-induced states rather than naturalistic attention, 2) report variance and per-subject results, and 3) perform a true cross-content generalization check (e.g., train on some videos, test on others). With those fixes, this could become a solid community resource.\n\nFor a referee: I'd send it out. The resource is valuable and the design flaw is addressable. I'd probably not cite it in its current form for attention-validity claims, but I'd keep an eye on a revised version.","headline":"MEMA is a useful public EEG dataset for online-learning attention, but the single-video-per-state design means the labels are as much about stimulus identity as about attention, and the paper needs a more careful validation story.","tokens_in":8660,"tokens_out":1980,"would_cite":false,"duration_ms":20033,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MEMA, a public multi-label EEG dataset for classifying mental attention states—neutral, relaxing, and concentrating—during online learning, and validates it through 1,060 minutes of recordings from 20 subjects…","keywords":["EEG dataset","attention classification","online learning","multi-label","mental attention states","electroencephalography","emotion labels","Big Five personality"],"falsifier":"A direct test would be to rerun the same paradigm with an independent objective measure of attention—for instance, recording eye-gaze patterns, reaction times to intermittent probes, or a second EEG-based attention index—and compare those measures across the neutral, relaxing, and concentrating conditions. If the relaxing and concentrating conditions produce indistinguishable objective attention, or if self-report labels frequently disagree with the objective measure, the dataset's core labeling assumption would fail.","tokens_in":7767,"feed_emoji":"🧠","tokens_out":8668,"duration_ms":78623,"temperature":0.7,"pith_summary":"This paper presents MEMA, a publicly released multi-label electroencephalography (EEG) dataset for classifying three mental attention states during online learning: neutral, relaxing, and concentrating. The data come from 20 subjects who each completed 12 randomized trials, yielding 1,060 minutes of EEG recordings alongside self-reported emotion labels, personal information, and Big Five personality traits. The paper argues that the field needs such a public resource because existing EEG attention datasets are scarce, often non-public, and collected under inconsistent paradigms. It validates the dataset by showing expected alpha and beta power patterns, baseline classification accuracies up to 85.12% subject-dependent and 64.84% cross-subject, and a statistical link between attention and emotional valence and arousal.","feed_headline":"Public EEG dataset tags attention states in online learning","feed_subtitle":"MEMA logs 1,060 minutes of brain signals from 20 learners and links attention to emotion and personality.","key_machinery":"The central object is the MEMA dataset together with its standardized three-task collection paradigm. Each attention state is induced by a specific video—neutral, relaxing, and concentrating—and each trial ends with self-assessment of attention and emotion, so the labels are self-reports anchored to a controlled stimulus. Around this, the validation machinery includes EEG preprocessing (notch filtering, band-pass filtering, ICA artifact removal), topographic analysis of $\\alpha$ and $\\beta$ power over frontal sites, six baseline classifiers evaluated in subject-dependent and cross-subject settings, a $\\chi^2$ test for the attention-emotion association, and hard parameter sharing multi-task learning to test which emotion dimension best pairs with attention classification.","core_discovery":"The central claim is that MEMA is a validated, publicly available multi-label EEG dataset for mental attention state classification in online learning, with three classes—neutral, relaxing, and concentrating—and auxiliary labels that make it more than a single-task collection. The collection paradigm assigns a distinct video task to each state: a blank-screen neutral clip, a soothing scenic clip with music for relaxing, and a machine-learning lecture clip for concentrating, followed by a comprehension question. After each trial, subjects self-report their attention state and rate emotion on the valence-arousal-dominance model. The paper demonstrates that the dataset supports reliable attention classification with classical and deep learning models, that $\\alpha$ power rises and $\\beta$ power falls as attention decreases, and that attention is statistically associated with valence and arousal; in particular, multi-task learning that pairs attention with valence improves classification accuracy.","pith_inferences":["A natural extension would be to use MEMA to train real-time attention monitors for lecture-style video content, since the concentrating task uses a genuine course video and the labels are trial-level.","Because the attention labels are self-reports, an independent behavioral or physiological check—such as eye tracking or response-time measures during the concentrating task—could strengthen the validity of the ground truth, something the paper itself does not report.","With the included personality and personal-information fields, the dataset could support individual-difference analyses, for example whether personality traits modulate the strength of the alpha-beta attention signature, but such analyses are not carried out in this paper."],"forward_implications":["Researchers gain a public, multi-label benchmark for EEG-based attention classification in online learning, addressing the scarcity that has limited reproducibility and comparability.","The presence of emotion labels, personality traits, and personal information enables studies of how attention interacts with affect and individual differences, not just single-label classification.","The standardized paradigm—three states, randomized trial order, fixed durations informed by physiological and psychological research—offers a template for future EEG data collections in learning contexts.","The reported baselines (up to 85.12% subject-dependent and 64.84% cross-subject accuracy) provide concrete reference points for evaluating new attention classifiers.","Because valence is the emotion dimension statistically tied to attention, the paper's multi-task results imply that attention and valence should be modeled together, not paired arbitrarily with arousal or dominance."],"supporting_citations":[{"why":"Supplies the prior online-learning attention EEG collection that is non-public, establishing the scarcity gap MEMA addresses.","marker":"[6]"},{"why":"Provides a comparison dataset in auditory attention with a non-public collection that informs the paradigm design.","marker":"[7]"},{"why":"Comparison dataset used in Table I, showing a public but single-subject attention EEG collection.","marker":"[8]"},{"why":"Prior attention detection study whose task design and non-public dataset the authors build on.","marker":"[9]"},{"why":"Supplies the claim that attention is intertwined with emotion and that valence is the emotion dimension most related to attention.","marker":"[10]"},{"why":"Comparison dataset in Table I that is non-public and includes eye-gaze data, used to position MEMA's public multi-label coverage.","marker":"[16]"},{"why":"Comparison dataset in Table I from e-learning attention feedback, used to position MEMA's coverage.","marker":"[17]"},{"why":"The public online-learning EEG dataset MEMA extends with more trials and auxiliary labels.","marker":"[18]"},{"why":"Defines the valence-arousal-dominance model used for the emotion self-report labels.","marker":"[25]"},{"why":"Provides the Chi-squared test of independence used in the multi-label correlation study.","marker":"[36]"}],"fun_headline_variants":["EEG dataset captures attention states in online learning","New EEG dataset links attention to emotion in e-learning","MEMA: 1,060 minutes of EEG for attention classification","Multi-label EEG set for studying attention in online classes","EEG attention data from 20 learners with emotion labels"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three video tasks reliably induce the intended attention states and that subjects' self-reported attention labels accurately capture those states; if the videos fail to induce the intended states, or if self-reports are inaccurate, then the classification results and attention-emotion correlations describe the videos or the reports rather than actual brain-state attention.","fun_headline_variants_meta":{"raw":{"variants":["EEG dataset captures attention states in online learning","New EEG dataset links attention to emotion in e-learning","MEMA: 1,060 minutes of EEG for attention classification","Multi-label EEG set for studying attention in online classes","EEG attention data from 20 learners with emotion labels"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000687,"raw_usage":{"total_tokens":3119,"prompt_tokens":957,"completion_tokens":2162,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":573,"tokens_out":2162,"duration_ms":15691,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T20:11:34.686481+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would be to rerun the same paradigm with an independent objective measure of attention—for instance, recording eye-gaze patterns, reaction times to intermittent probes, or a second EEG-based attention index—and compare those measures across the neutral, relaxing, and concentrating conditions. If the relaxing and concentrating conditions produce indistinguishable objective attention, or if self-report labels frequently disagree with the objective measure, the dataset's core labeling assumption would fail.","supporting_citations":[{"cited_title":"Attention recognition system in online learning platform using eeg signals,","cited_arxiv_id":null,"evidence_quote":"Supplies the prior online-learning attention EEG collection that is non-public, establishing the scarcity gap MEMA addresses."},{"cited_title":"Detection of sustained auditory attention in students with visual impairment,","cited_arxiv_id":null,"evidence_quote":"Provides a comparison dataset in auditory attention with a non-public collection that informs the paradigm design."},{"cited_title":"Eeg-based closed-loop neurofeed- back for attention monitoring and training in young adults,","cited_arxiv_id":null,"evidence_quote":"Comparison dataset used in Table I, showing a public but single-subject attention EEG collection."},{"cited_title":"Detection of human attention using eeg signals,","cited_arxiv_id":null,"evidence_quote":"Prior attention detection study whose task design and non-public dataset the authors build on."},{"cited_title":"Attention recognition in eeg- based affective learning research using cfs+ knn algorithm,","cited_arxiv_id":null,"evidence_quote":"Supplies the claim that attention is intertwined with emotion and that valence is the emotion dimension most related to attention."},{"cited_title":"Electroencephalogram-based attention level classification using convolution attention memory neural network,","cited_arxiv_id":null,"evidence_quote":"Comparison dataset in Table I that is non-public and includes eye-gaze data, used to position MEMA's public multi-label coverage."},{"cited_title":"Eeg-based attention feedback to improve focus in e-learning,","cited_arxiv_id":null,"evidence_quote":"Comparison dataset in Table I from e-learning attention feedback, used to position MEMA's coverage."},{"cited_title":"The eeg-based attention analysis in multimedia m-learning,","cited_arxiv_id":null,"evidence_quote":"The public online-learning EEG dataset MEMA extends with more trials and auxiliary labels."},{"cited_title":"Three dimensions of emotion","cited_arxiv_id":null,"evidence_quote":"Defines the valence-arousal-dominance model used for the emotion self-report labels."}],"review_version":1}