{"id":"c993ebd4-b7cb-497b-b1ef-6894dd1f9e6d","arxiv_id":"2412.01587","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Handwriting strokes from both hands, classified by a CNN, yield degree-of-handedness scores that correlate with Edinburgh Inventory scores in a 43-person pilot.","lead":"This pilot study automatically grades degree of handedness from handwriting signals, using stroke features and machine learning to tell dominant from non-dominant hand writing. If validated, it could offer a quantitative replacement for handedness questionnaires in behavioral, clinical, and forensic work.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 4-point DoH score is a linear rescaling of dominant-vs-non-dominant classification accuracy, but the paper never tests the core hypothesis that this accuracy is a monotone surrogate for degree of handedness; group-level ordering against self-identified categories is not enough.","rationale":"Good-faith reading: the leave-one-subject-out results are real and the U>PU>A ordering is exactly what a handedness measure should look like, so this is not an accusation of fabrication. However, the leap from 'dominant vs non-dominant classification accuracy' to 'degree of handedness score' is the load-bearing step, and the paper explicitly labels it a hypothesis rather than a tested result. The reader's weakest_assumption captures the same issue. Since the paper is a pilot with no code/data and a small, partly self-selected sample, this unvalidated surrogate assumption is what makes the verdict CONDITIONAL rather than ACCEPT. It does not require REJECT because the objective kinematic signal and LOSO ordering are plausible and independently checkable; a failed validation would raise the concern to a rejection criterion. Thus the reader's verdict stands unchanged.","tokens_in":17424,"tokens_out":8354,"duration_ms":84081,"concrete_test":"On a new or released cohort, collect for each subject: (1) LOSO CNN D/ND accuracy from the same protocol, (2) an objective quantitative laterality measure such as Annett pegboard quotient or finger-tapping asymmetry, and (3) Edinburgh Inventory plus self-reported practice hours. Pre-specify the accuracy-to-4-point mapping and all analysis before running it. Compute Spearman rank correlation between CNN accuracy and the quantitative laterality quotient, and the partial correlation controlling for practice hours. If the correlation is weak or the practice-adjusted correlation vanishes, the monotone-surrogate assumption is falsified; if it survives, the central grading claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 4-point DoH score in Section II.H Part C is a direct linear rescaling of the leave-one-subject-out dominant-vs-non-dominant classification accuracy (100→4, 0→0). The paper states the accuracy-to-DoH relationship as a hypothesis ('It was hypothesized...') but never independently validates it. High per-subject accuracy only shows that the CNN can separate that subject's dominant and non-dominant stroke kinematics; this separability can reflect practice, task familiarity, motivation, trial-order effects, or idiosyncratic stroke patterns rather than an underlying 'inherent capability' of the brain. The observed U>PU>A ordering is close to definitional because the PU group was defined by non-dominant-hand training and the A group by equal writing fluency. The only external check, agreement with the Edinburgh Inventory, is weakened by the fact that the same EI/self-report information was used to form the groups, and the non-linear curves in Eqs. (2)-(6) are fit in-sample before correlations and Bland-Altman agreement are computed. So the monotone-surrogate assumption is the load-bearing but untested link between 'the CNN can tell the hands apart' and 'the CNN grades degree of handedness.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes automated grading of degree of handedness (DoH) from digitizer-captured handwriting. Forty-three subjects were grouped as Unidextrous (U), Partially-Unidextrous (PU), or Ambidextrous (A). Stroke-level static and dynamic features were used with three approaches: a Davies-Bouldin Index score, an MLP classifier, and a CNN classifier. The CNN reaches 95.06±3.08% accuracy in stratified 10-fold cross-validation for dominant-vs-non-dominant hand classification. Leave-one-subject-out accuracies are converted through a linear mapping into a 4-point DoH score. The computed scores are compared with the Edinburgh Inventory using Bland-Altman plots, correlation, and RMSE, with roughly 90% of scores reported within the 95% confidence interval. The authors conclude that handwriting can provide a more resolved, quantitative DoH assessment than questionnaires.","tokens_in":17686,"tokens_out":3336,"duration_ms":32934,"significance":"If the central mapping from classification accuracy to degree of handedness were independently validated, this would be a valuable, low-cost, quantitative supplement to the Edinburgh Inventory, with plausible applications in neurorehabilitation, psychometry, and forensics. The paper includes strengths that should be acknowledged: a real data-collection protocol with ethics approval, multiple baseline methods (DB index, six classical ML classifiers, MLP, CNN), ablation studies for the MLP and CNN architectures, per-subject leave-one-subject-out results in Table II, and an explicit limitations paragraph. The pilot nature and small sample are acknowledged. However, two load-bearing issues prevent the results from currently supporting the paper's central claim: the monotone-surrogate assumption connecting per-subject classification accuracy to DoH is never independently tested, and the agreement with the Edinburgh Inventory is computed after in-sample curve fitting, which inflates agreement. The headline 95% accuracy is also based on stroke-level cross-validation with subject leakage.","major_comments":[{"comment":"The 4-point DoH score is defined as a linear rescaling of the leave-one-subject-out dominant-vs-non-dominant classification accuracy (100 mapped to 4, 0 mapped to 0). The paper states as a hypothesis that high accuracy implies unidexterity and low accuracy implies ambidexterity, but this monotone-surrogate assumption is never independently validated. High per-subject classification accuracy only shows that the model can separate that subject's dominant and non-dominant stroke kinematics; this separability may reflect practice, task familiarity, trial-order effects, or idiosyncratic stroke patterns rather than an inherent degree of handedness. The observed U>PU>A ordering is also close to definitional, because the PU group was defined by non-dominant-hand training and the A group by equal writing fluency. I request additional evidence for monotonicity, for example by correlating per-subject accuracies with a quantitative, independently measured hand-performance asymmetry that is not used to form the groups, or by showing that accuracy predicts EI scores in a held-out split with no in-sample curve fitting.","section":"Section II.H Part C"},{"comment":"The abstract and Section III report 'average classification accuracy of 95.06±3.08%' from stratified 10-fold cross-validation. Because the cross-validation folds are created at the stroke level, strokes from the same subject appear in both training and test folds, violating the independence assumption and inflating accuracy. The subject-level generalization results are the leave-one-subject-out accuracies in Table II, which are considerably lower and more variable across subjects. The paper should report the stroke-level CV as a within-subject discrimination result and make the leave-one-subject-out results the primary basis for the DoH-grading claims, or alternatively use a nested or grouped cross-validation that respects subject boundaries when reporting the headline accuracy.","section":"Section III.A Part B / Table II"},{"comment":"The comparison with the Edinburgh Inventory is compromised by in-sample curve fitting. The non-linear scaling functions for EI scores (Eqs. 2, 5, 6) are obtained by fitting curves to the same subjects whose scores are later used to compute correlations, Bland-Altman agreement, and RMSE. This procedure cannot provide an unbiased estimate of agreement; for example, a sufficiently flexible in-sample curve can make almost any two monotone sequences agree. I recommend either reporting the unscaled comparison, or using a cross-validated calibration in which scaling parameters are estimated on a training subset and evaluated on held-out subjects. Without this, the statement that 'around 90% of the obtained scores... were found to be in accordance with the EI scores under 95% confidence interval' is not supported.","section":"Section I, Eqs. (2)-(6)"},{"comment":"The ambidextrous group contains only two subjects, and one of them (S22) has a condition that limits use of one hand for long periods, making the participant's handwriting data potentially atypical. The claim that the method can grade ambidexterity rests almost entirely on these two data points. I do not ask for a larger cohort in this pilot, but the conclusions should explicitly state that the ambidextrous end of the spectrum is not empirically established; the current text already mentions sample-size limitations, and I would like the Discussion and Conclusion to make this limitation more prominent when claiming that the method differentiates all three DoH categories.","section":"Section II.B / Table II"}],"minor_comments":[{"comment":"The abstract states '95.06%' without the ±3.08% standard deviation; the full result appears in Section III, but the abstract should at least mention that this is a stroke-level stratified CV result rather than a subject-level generalization.","section":"Abstract / Section III.A Part B"},{"comment":"The introduction cites the Edinburgh Inventory [18] but does not mention Annett's questionnaire [34] in the main text until Section II.B; consider citing [34] in the introduction where questionnaire-based methods are listed.","section":"Section I (Introduction) / References"},{"comment":"In Fig. 4, the panels {a}, {b}, and {c} are referenced in the text but the figure caption does not explain the visual elements in each panel; please expand the caption so that a reader can understand the network diagrams and the DB-index illustration without returning to the text.","section":"Fig. 4"},{"comment":"The feature list in Table I includes 'PV' and 'Number of Strokes' but the supplementary description for 'PV' appears as 'The product of average absolute velocity and average pen pressure per segment'; please align the notation and define 'PV' in the table caption.","section":"Section II.F / Table I"},{"comment":"The sentence 'It is interesting to note that the EI scores' resolution were limited' contains a subject-verb agreement issue; please revise to 'the EI scores' resolution was limited'.","section":"Section III.B"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a worthwhile and under-studied problem, and the authors are transparent about several limitations. However, the two main statistical issues—subject leakage in the headline CV accuracy and in-sample nonlinear scaling before agreement testing—are load-bearing for the central claim. I think the work can be revised to be publishable, but it needs a substantial methodological overhaul: subject-level cross-validation for all reported metrics, an independent or cross-validated validation of the accuracy-to-DoH monotonicity assumption, and a clear separation of in-sample curve fitting from out-of-sample agreement assessment. I would also encourage a statistician to review the Bland-Altman and RMSE calculations, because the current reporting mixes in-sample and out-of-sample evidence. The small number of ambidextrous subjects (n=2) further limits the generality, and I would expect the revision to temper the conclusions accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the 95.06% accuracy in the abstract is from stroke-level 10-fold CV, so the same subject's strokes are in both training and test folds. That number is not evidence for person-level handedness grading. Second, the 4-point DoH score is just a linear rescaling of per-subject classification accuracy, and the paper never independently tests the hypothesis that accuracy is a monotone surrogate for degree of handedness. The group ordering (U > PU > A) is consistent, but those groups are defined by self-reported fluency and practice, so that ordering is about as circular as it can be without being tautological.\n\nThat said, the paper is not empty. The task—quantifying degree of handedness on a continuum from handwriting, rather than classifying direction—is genuinely new in the cited literature. Using segmented strokes (zero crossings of vertical velocity) and letting a CNN operate on raw kinematic traces is a reasonable extension of existing handwriting-analysis methods. And the leave-one-subject-out accuracies, which are the ones actually used for grading, do show a clear ordering across the three self-identified groups, with only modest overlap. That is the strongest evidence in the paper, and it suggests the signal is real, even if the pilot is too small and too confounded to prove it.\n\nThe soft spots are proportionate to a pilot. Only two ambidextrous subjects, no code or data release, and several architecture details left heuristic. The comparison to the Edinburgh Inventory is weakened because the non-linear curves in Eqs. (2), (5), and (6) are fit in-sample before correlations and Bland-Altman agreement are reported. The stress-test note is right that this is the load-bearing link: accuracy can reflect practice, familiarity, motivation, or idiosyncratic motor patterns rather than an inherent 'degree of handedness.' The paper does not address that.\n\nWho gets value? Researchers in laterality and motor control will want to know the approach exists, but should not treat the graded scores as validated. The citation pattern is fine—standard references in handedness and handwriting analysis. This is a pilot that identifies a promising direction and a set of pitfalls, not a finished instrument. I'd send it to a serious referee, but with instructions to demand subject-disjoint validation as the headline result, pre-specified accuracy-to-score mapping, and independent validation or at least a much larger cohort.","headline":"The pilot's headline accuracy is inflated by subject leakage and its grading scale rests on an unvalidated monotonicity assumption, but the subject-disjoint ordering of accuracies makes the underlying question worth a serious referee.","tokens_in":18232,"tokens_out":4418,"would_cite":false,"duration_ms":36729,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that degree of handedness can be graded automatically from handwriting strokes, with a CNN separating dominant from non-dominant strokes at 95.06% accuracy and computational scores agreeing with the Edinburgh Inventory…","keywords":["degree of handedness","handwriting analysis","handwriting strokes","convolutional neural network","Edinburgh Handedness Inventory","Davies-Bouldin index","lateralization","automated grading"],"falsifier":"Train the CNN with leave-one-subject-out on the same 43-subject dataset, then take a new group of strongly unidextrous subjects and have them write the same tasks twice: once with their non-dominant hand completely unpracticed and once after a short, controlled training session with that hand. If per-subject classification accuracy, and therefore the 4-point score, drops substantially after training while the Edinburgh Inventory score and self-reported preference are unchanged, the measure is tracking differential skill rather than an inherent degree of handedness.","tokens_in":17164,"feed_emoji":"✍️","tokens_out":6704,"duration_ms":56069,"temperature":0.7,"pith_summary":"This pilot study tries to establish that degree of handedness—how strongly someone favors one hand, as opposed to merely which hand they prefer—is measurable from handwriting. The authors segment writing into strokes using zero crossings of vertical velocity, extract kinematic and static features, and grade each subject by how well a classifier can separate their dominant-hand strokes from their non-dominant-hand strokes. A convolutional neural network operating directly on stroke coordinates and timestamps reached 95.06±3.08% average accuracy in dominant/non-dominant classification under stratified 10-fold cross-validation. Converted to a 4-point score, the CNN and MLP grades agreed with Edinburgh Inventory scores for roughly 90% of subjects within a 95% confidence interval, while giving finer resolution than the questionnaire. If right, this gives a cheap, quantitative, continuous alternative to questionnaire-based handedness assessment for rehabilitation, brain-computer interfaces, psychometry, and forensics.","feed_headline":"CNN reads degree of handedness from handwriting strokes","feed_subtitle":"A 43-person pilot study matches a 4-point handwriting score to Edinburgh Inventory scores for about 90% of subjects.","key_machinery":"The load-bearing mechanism is the handwriting stroke: each writing trial is segmented between successive zero crossings of vertical velocity, so every stroke is one elementary movement unit. For the statistical and MLP pipelines, 25 time, static, and dynamic features are extracted per stroke; the DB index sums cluster-separation scores across features for each subject's dominant vs non-dominant strokes. The CNN bypasses manual features, taking padded x, y, and time channels of each stroke through three 1D convolutional layers (128, 64, 32 filters) and two fully connected layers. The grading assumption connects all of these: a high dominant-vs-nondominant classification accuracy means the two hands produce distinctly different stroke statistics, which is interpreted as a high degree of handedness, and that accuracy is linearly rescaled into a 4-point score where 4 means strongly unidextrous and 0 means fully ambidextrous.","core_discovery":"On its own terms, the paper's central claim is that the separability of a person's dominant-hand and non-dominant-hand handwriting strokes is a valid proxy for their degree of handedness. Subjects whose strokes form two well-separated classes are \"unidextrous\"; subjects whose stroke distributions overlap are \"ambidextrous\"; intermediate separability is \"partially unidextrous\". The paper operationalizes this by training a CNN on raw x, y, timestamp stroke channels with leave-one-subject-out evaluation: the test subject's classification accuracy is linearly mapped to a 0–4 degree-of-handedness score. The CNN achieved 95.06±3.08% classification accuracy under stratified 10-fold cross-validation, and 90.6% of the computational scores (DB, MLP, CNN) fell within the 95% confidence interval of the Edinburgh Inventory scores after nonlinear scaling, with correlations of 0.87 (DB) and 0.91 (MLP and CNN) against scaled EI. The authors further claim this approach resolves differences the Edinburgh Inventory cannot see, since subjects with identical inventory scores received distinct computational scores.","pith_inferences":["The 0–4 scale is obtained by linearly mapping classification accuracy to a score; the paper does not establish that the resulting scale has equal intervals, so \"one point\" differences may not be comparable across the range. A reader should treat the score as ordinal until a calibration study is done.","If the separability measure is really tracking degree of handedness, the same stroke-separability logic should transfer to other fine-motor activities such as drawing, tracing, or tapping, and could be tested as a forensic screen for feigned handedness, where subjects may control which hand they use but have trouble controlling stroke dynamics.","A direct testable extension: measure a subject's CNN score before and after a short block of non-dominant-hand writing practice. If the score shifts toward ambidexterity with practice while the Edinburgh score and self-reported preference stay fixed, the measure is partly a skill-asymmetry metric rather than a fixed trait."],"forward_implications":["Degree of handedness can be scored from a single short digitizer session, without expensive neuroimaging, giving a continuous quantitative measure where the Edinburgh Inventory gives only coarse semi-quantitative categories.","Because the CNN also labels strokes as dominant or non-dominant, a single model yields both the degree and the direction of handedness, whereas the DB-index method yields only degree.","Task selection matters: the loop-writing task (Task 7) separated hands best and the \"llllll\" task (Task 1) worst, so future versions could drop or replace low-information tasks.","The gender difference in scores (significant, p<<0.05) and the absence of a left/right difference suggest the measure is sensitive to lateralization-related motor organization, not merely to which hand is reported as dominant."],"supporting_citations":[{"why":"Supplies the Edinburgh Inventory questionnaire baseline whose scores are compared against the computational grades.","marker":"[18]"},{"why":"Identifies handwriting as the best motor test for handedness, justifying the choice of writing tasks.","marker":"[22]"},{"why":"Provides the vertical-velocity oscillation model used for segmenting handwriting into strokes.","marker":"[37]"},{"why":"Motivates the digital recording and preprocessing methods, including filtering of handwriting movements.","marker":"[36]"},{"why":"Supports the use of nonsensical cursive words as writing tasks for handedness assessment.","marker":"[35]"},{"why":"Supplies the handwriting-feature and task paradigm adapted from Parkinson's disease assessment.","marker":"[27]"},{"why":"Contributes to the preliminary hand-preference questionnaire used to form the three subject groups.","marker":"[34]"}],"fun_headline_variants":["CNN grades handedness from handwriting strokes","Handwriting strokes reveal degree of handedness","AI matches self-reported handedness in 90% of pilot study","New CNN method scores handedness from pen strokes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that the ease with which a classifier separates a person's dominant-hand strokes from their non-dominant-hand strokes is a faithful, monotone measure of that person's degree of handedness, rather than a reflection of practice, task familiarity, motivation, or idiosyncratic stroke style.","fun_headline_variants_meta":{"raw":{"variants":["CNN grades handedness from handwriting strokes","Handwriting strokes reveal degree of handedness","AI matches self-reported handedness in 90% of pilot study","New CNN method scores handedness from pen strokes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1469,"prompt_tokens":1043,"completion_tokens":426,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":366}},"tokens_in":659,"tokens_out":426,"duration_ms":3933,"temperature":1.0,"reasoning_tokens":366,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:17:14.174873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the CNN with leave-one-subject-out on the same 43-subject dataset, then take a new group of strongly unidextrous subjects and have them write the same tasks twice: once with their non-dominant hand completely unpracticed and once after a short, controlled training session with that hand. If per-subject classification accuracy, and therefore the 4-point score, drops substantially after training while the Edinburgh Inventory score and self-reported preference are unchanged, the measure is tracking differential skill rather than an inherent degree of handedness.","supporting_citations":[{"cited_title":"Cerebral asymmetry and lan guage development: Cause, correlate, or consequence?,","cited_arxiv_id":null,"evidence_quote":"Supplies the Edinburgh Inventory questionnaire baseline whose scores are compared against the computational grades."},{"cited_title":"Unidextrous","cited_arxiv_id":null,"evidence_quote":"Identifies handwriting as the best motor test for handedness, justifying the choice of writing tasks."},{"cited_title":"EMOTHAW: A Novel Database for Emotional State Recognition from Handwriting and Drawing,","cited_arxiv_id":null,"evidence_quote":"Provides the vertical-velocity oscillation model used for segmenting handwriting into strokes."},{"cited_title":"The reliability of some motor performance tests of handedness,","cited_arxiv_id":null,"evidence_quote":"Supports the use of nonsensical cursive words as writing tasks for handedness assessment."},{"cited_title":"Why are consistently -handed individuals more authoritarian? The role of need for cognitive closure,","cited_arxiv_id":null,"evidence_quote":"Supplies the handwriting-feature and task paradigm adapted from Parkinson's disease assessment."},{"cited_title":"Purdue Pegboard Test,","cited_arxiv_id":null,"evidence_quote":"Contributes to the preliminary hand-preference questionnaire used to form the three subject groups."}],"review_version":1}