{"id":"b73678cc-646b-4ec0-9b32-683433fa8b1c","arxiv_id":"2605.12408","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FAAR is a new automated artifact rejection method using compact features and adaptive Signal Quality Index thresholds that improves MI-BCI performance most in low-baseline conditions and reduces inter-subject variability across 13 public datasets.","lead":"The paper presents FAAR, an automated lightweight method for rejecting EEG artifacts in motor imagery BCIs that shows subject- and regime-dependent benefits, especially in low-SNR settings, while reducing inter-subject variability. A smart generalist might read it to see how data curation choices affect real-world BCI reliability beyond decoder design alone.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flags the generalization of the adaptive SQI mechanism as the critical assumption given abstract-only access. With full text now available, the multi-dataset empirical results directly address that assumption, leaving no additional load-bearing gap that would alter the UNVERDICTED verdict.","tokens_in":1730,"tokens_out":280,"duration_ms":24181,"concrete_test":"Re-run the 13-dataset comparison using only the no-rejection baseline and FAAR; confirm that the reported subject-wise accuracy distributions show reduced variance (e.g., via Levene test p<0.05) specifically in the low-SNR subset while preserving the regime-dependent pattern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that rejection effects are strongly subject- and regime-dependent with largest gains in low-baseline/low-SNR conditions, and that FAAR reduces inter-subject variability without aggressive removal—rests on evaluation across 13 public MI datasets with comparisons to baselines. The reader's weakest assumption (reliable artifact identification via compact features, SQI, and adaptive thresholding without prior knowledge or manual tuning) is the key precondition, but the abstract plus full-text evaluation description provides direct empirical support via the multi-dataset results. No internal inconsistency or untested assumption that would falsify the headline claims is apparent.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Fast Automatic Artifact Rejection (FAAR), a lightweight automated method for EEG artifact rejection in motor imagery (MI) BCIs. FAAR extracts a compact set of artifact-sensitive features, computes an epoch-level Signal Quality Index, and applies adaptive threshold selection to reject contaminated epochs without prior artifact knowledge or manual tuning. It evaluates the approach on 13 public MI datasets, comparing against a no-rejection baseline, AutoReject, and Isolation Forest, and reports that rejection effects are strongly subject- and regime-dependent (largest gains in low-baseline/low-SNR conditions), that FAAR reduces inter-subject performance variability without aggressive data removal, and that the method is consistent across offline, training, and online regimes while satisfying real-time constraints.","tokens_in":1843,"tokens_out":391,"duration_ms":20897,"significance":"If the multi-dataset empirical results hold, the work provides concrete evidence that automated artifact rejection should be treated as an adaptive, regime-dependent component of MI-BCI pipelines rather than a fixed preprocessing step. The finding that gains are largest under low-SNR conditions and that inter-subject variability is reduced without heavy data loss directly addresses BCI illiteracy and reliability issues; the lightweight, fully automated design also supports deployment under real-time constraints.","major_comments":[],"minor_comments":[{"comment":"Abstract: states that evaluation results and comparisons were performed but supplies no quantitative performance numbers, statistical tests, or subject-exclusion criteria, which weakens the reader's ability to gauge the magnitude of the reported subject-dependent gains and variability reduction from the abstract alone.","section":"Abstract"},{"comment":"The description of the compact artifact-sensitive feature set and the derived Signal Quality Index would benefit from an explicit enumeration or pseudocode in the methods section to allow exact reproduction.","section":"Methods"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive summary and recommendation of minor revision. The assessment correctly captures the core contributions of FAAR as a lightweight, adaptive artifact rejection method whose benefits are regime- and subject-dependent, and we appreciate the recognition that these results speak to BCI reliability and illiteracy issues.","responses":[],"tokens_in":1298,"tokens_out":78,"duration_ms":16071,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"FAAR combines a compact feature set, an epoch-level signal quality index, and per-subject adaptive thresholding into a fully automatic pipeline. The paper runs it on 13 public motor-imagery datasets against a no-rejection baseline, AutoReject, and Isolation Forest. The central result is that rejection effects are strongly subject- and regime-dependent, with the largest improvements in low-baseline, low-SNR conditions, and that the method shrinks inter-subject performance spread without aggressive data loss. That pattern matters for BCI reliability and the BCI-illiteracy problem. The work also checks that the same pipeline behaves consistently from offline curation through online filtering, which is a real constraint for actual systems. The evaluation is broad enough that the regime-dependent pattern is unlikely to be a single-dataset artifact. The absence of manual tuning and prior artifact-type knowledge is a clear practical plus over many existing tools. The main soft spot is that the abstract gives no effect sizes, confidence intervals, or statistical tests, so the strength of the variability-reduction claim has to be judged from the full tables and figures. If those numbers turn out modest or if the chosen features miss certain artifact classes on new hardware, the adaptive advantage could shrink. The feature set itself is not derived from first principles, so its generality rests on the empirical coverage. This paper is aimed at people who build or maintain end-to-end BCI pipelines rather than pure decoder theorists. Anyone who has to ship a system that works across subjects will find the regime-dependent results and the low tuning burden useful. It is grounded enough in public data and clear baselines to deserve a serious referee, even if the core ideas are incremental combinations of existing techniques.","headline":"FAAR gives a practical, low-overhead way to do adaptive artifact rejection in MI-BCIs and the 13-dataset results back the claim that gains are biggest where baseline SNR is poor.","tokens_in":2346,"tokens_out":420,"would_cite":false,"duration_ms":23252,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical BCI artifact-rejection pipeline (FAAR + SQI + knee thresholding) has no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"Paper's central machinery (compact artifact features, epoch-level SQI, self-calibrated reference, adaptive knee detection) is a practical signal-processing heuristic evaluated on 13 MI datasets. RS framework derives J-cost, φ, 8-tick periodicity, 3D spacetime and constants from a single distinction (AbsoluteFloorClosure, Cost/FunctionalEquation, DimensionForcing, AlexanderDuality). No ratio-symmetric cost, golden-ratio ladder, or parameter-free derivation appears; domain is applied neuroscience/ML, where RS expresses no opinion.","tokens_in":46243,"confidence":"high","tokens_out":164,"duration_ms":12083,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Automated artifact rejection improves motor imagery BCI decoding most for low-baseline subjects and reduces performance spread across users.","keywords":["motor imagery","BCI","artifact rejection","EEG","signal quality","decoding performance","FAAR"],"falsifier":"A new MI dataset in which applying FAAR to low-baseline subjects produces no accuracy gain or increases the spread of performance across subjects.","tokens_in":2663,"feed_emoji":"🧠","tokens_out":488,"duration_ms":34185,"temperature":0.7,"pith_summary":"The paper proposes FAAR as a lightweight automated method to reject contaminated EEG epochs in motor imagery tasks by building a Signal Quality Index from a small set of artifact-sensitive features and choosing thresholds without manual input. Tests across 13 public datasets show that the benefits of rejection depend on the individual subject and the starting signal quality, delivering the clearest gains when baseline accuracy or SNR is already low. The approach also narrows the gap in results between different subjects without discarding large amounts of data, which matters for making BCIs more reliable for a wider range of users.","feed_headline":"Artifact rejection boosts low-SNR motor imagery BCIs most","feed_subtitle":"FAAR cuts inter-subject variability across 13 datasets without heavy data loss and works in real time.","key_machinery":"Fast Automatic Artifact Rejection (FAAR), a method that builds an epoch-level Signal Quality Index from artifact-sensitive features and applies adaptive thresholding to identify and reject contaminated epochs.","core_discovery":"FAAR computes a compact set of artifact-sensitive features, derives an epoch-level Signal Quality Index, and adaptively selects rejection thresholds to remove contaminated epochs without prior knowledge of artifact types or manual tuning. Evaluated on 13 MI datasets against a no-rejection baseline, AutoReject, and Isolation Forest, the method produces subject- and regime-dependent effects on decoding accuracy, with the largest improvements in low-baseline or low-SNR conditions, while reducing inter-subject performance variability without aggressive data removal and maintaining consistent behavior across offline, training, and online settings.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Artifact rejection effects vary by subject and SNR in MI BCIs","FAAR shows adaptive rejection impact on 13 MI datasets","Inter-subject MI decoding variability reduced via FAAR","Regime-dependent rejection outcomes in motor imagery BCIs","FAAR yields consistent artifact handling across BCI regimes"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"A compact set of artifact-sensitive features and the derived Signal Quality Index, together with adaptive threshold selection, can reliably flag contaminated epochs across many different MI datasets without knowing the artifact types ahead of time or needing manual tuning.","fun_headline_variants_meta":{"raw":{"variants":["Artifact rejection effects vary by subject and SNR in MI BCIs","FAAR shows adaptive rejection impact on 13 MI datasets","Inter-subject MI decoding variability reduced via FAAR","Regime-dependent rejection outcomes in motor imagery BCIs","FAAR yields consistent artifact handling across BCI regimes"]},"model":"grok-4.3","cost_usd":0.005202,"raw_usage":{"total_tokens":2457,"prompt_tokens":699,"num_sources_used":0,"completion_tokens":76,"cost_in_usd_ticks":52015500,"prompt_tokens_details":{"text_tokens":699,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1682,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":699,"tokens_out":76,"duration_ms":65257,"temperature":1.0,"reasoning_tokens":1682,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T06:13:39.421338+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new MI dataset in which applying FAAR to low-baseline subjects produces no accuracy gain or increases the spread of performance across subjects.","supporting_citations":[],"review_version":2}