{"id":"1644bc0d-f9eb-4380-9044-b8c4bac7bbec","arxiv_id":"2607.25126","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Demographics and motivations correlate with OSS project-selection preferences, with distinct patterns for newcomers versus experienced practitioners in a 208-person survey.","lead":"A survey of 208 software practitioners finds that age, gender, region, experience, and OSS role correlate with why people join open-source projects, and those motives in turn correlate with preferred project traits like age, docs, and guidelines. Newcomers and experienced contributors differ systematically, which matters for onboarding design and project recommenders.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The correlation map rests on ~200+ Kruskal-Wallis tests with no evident family-wise correction; the paper's own tables (p=0.051, p=0.242 presented as findings) contradict the claim that \"all findings\" have Bonferroni-adjusted p<0.05.","rationale":"The reader's weakest assumption targeted external validity: convenience sampling, 66.6% screening attrition, sparse merged cells, and hypothetical self-reports. That is a fair concern, but it is one the paper partially concedes in §6, and it does not by itself invalidate sample-bound correlational claims. My concern is more load-bearing because it is internal: even taken purely as a description of these 208 respondents, many of the reported correlations likely do not survive the multiple-comparison burden the design creates, and the manuscript is internally inconsistent about its own correction procedure (the blanket \"all findings presented were statistically significant with Bonferroni adjusted p<0.05\" sentence cannot be true given p=0.051 and p=0.242 rows in the results tables). This is checkable, not speculative: the authors claim a full replication package, so the test family and corrected p-values can be recomputed exactly. I keep the verdict CONDITIONAL rather than REJECT because (a) several cells (Reputation×Region p=0.001, Networking×Region p=0.001, Career×Age p=0.003, Incentives×OSS-role p=0.001, plus the large-effect RQ2 cells like Learning×Guidelines ϵ²=0.232) plausibly survive proper correction, so a real core finding likely remains; (b) the RQ4 thematic analysis (κ=0.96/0.83) is unaffected; and (c) the fix is re-analysis and honest reframing, not new data collection. But the conditions should shift from the reader's framing-level fixes to a statistical one: re-run with family-wise correction, report the surviving cell set, drop or heavily caveat tiny-cell claims (gender \"Other\" n=5; Oceania n=12; pairwise post-hocs on cells of a handful of respondents), and address the response-style confound in RQ2 before the motivation→preference claims are used to justify recommender design.","tokens_in":26208,"tokens_out":2993,"duration_ms":90696,"concrete_test":"Using the Zenodo replication package, recompute every omnibus Kruskal-Wallis p-value for RQ1 (7 motivations × 7 demographics, per group) and RQ2 (7 motivations × 8 characteristics, per group), then apply a single Holm (or Bonferroni) correction within each RQ×group family. Tabulate which cells of Tables 4, 7, 10, 14, 16, 18 survive, and reconcile the p=0.051 (Table 10) and p=0.242 (Table 18) rows against the §4.1 claim that all presented findings have adjusted p<0.05. If only the strong cells survive (Reputation×Region, Networking×Region, Career×Age, Incentives×OSS-role), the headline claim must be narrowed from \"age, gender, region, education, SE experience, OSS role correlate with motivations\" to roughly region, age, and OSS role. As a robustness pass for RQ2, within-person-center (ipsatize) each respondent's seven motivation ratings and re-run one representative cell (Learning×Clear-G","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that demographics correlate with motivations (RQ1), motivations with project-characteristic preferences (RQ2), and demographics with preferences (RQ3). The statistical engine is Kruskal-Wallis omnibus tests run per motivation × demographic cell, separately for two groups. The test families are large: RQ1 is 7 motivations × 7 demographic factors × 2 groups ≈ 98 omnibus tests; RQ2 is 7 × 8 × 2 ≈ 112; RQ3 adds dozens more. The paper states results were \"considered statistically significant, with Bonferroni adjusted p-values less than 0.05\" (§4.1), but the reported tables appear to contain raw p-values: Table 4 lists Learning×Gender p=0.035, Learning×SE-experience p=0.049; Table 7 lists Incentives×SE-Experience p=0.054; Table 10 lists Learning×Multilingual-Documentation p=0.051; Table 18 lists Region×Project-Age p=0.242 — all inside tables of purportedly significant results, with effect sizes attached. A p of 0.242 is not significant under any correction, and p≈0.03–0.05 cells cannot survive Bonferroni across a ~49-test family (threshold ≈0.001). With ~49 tests per group in RQ1, ~2–3 false positives are expected by chance alone, which is the same order as the number of marginal cells driving claims like \"gender correlates with motivations.\" A secondary confound: RQ2 groups respondents by motivation intensity, but motivation ratings within a person are correlated and share a response-style (yea-saying) component; a general \"rates everything extremely\" factor would mechanically produce the monotone patterns in Tables 8–13 without any motivation-specific preference structure. Neither concern requires new data to check — the replication package suffices.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper reports an online survey of 208 software practitioners (85 \"newcomers to OSS\" with SE backgrounds but no OSS contributions, 123 experienced OSS practitioners), recruited via Prolific and LinkedIn and screened with Danilova-style programming questions (632 recorded → 208 analyzed). It examines correlations between demographics and seven OSS motivations (RQ1), motivations and eight project-characteristic preferences (RQ2), and demographics and preferences (RQ3), using Kruskal–Wallis omnibus tests with Bonferroni post-hoc correction and ε² effect sizes, analyzed separately for the two groups. RQ4 thematically analyzes an open-ended item on improving recommendation systems, yielding four themes (personal fit/growth, project characteristics, interactive/evolving recommenders, collaboration/team dynamics). The design is standard and mostly careful: ethics approval, pilot, mandatory items, dual-coder thematic analysis with reported κ, and a public replication package. However, the statistical reporting is internally inconsistent: the blanket claim that all reported findings have Bonferroni-adjusted p<0.05 is contradicted by the paper's own tables (p=0.051, 0.054, 0.242), and several reported \"significant\" pairwise differences of 1–5% are not credible under any correction at n=123. These issues are load-bearing for the RQ1–RQ3 claims.","tokens_in":26635,"tokens_out":5099,"duration_ms":39434,"significance":"If the results hold, the paper offers a useful comparative descriptive map of motivations and project-selection preferences across newcomer and experienced contributor groups — a split few prior surveys make explicitly — with direct implications for OSS recommender design and maintainer practice. Strengths worth naming: a public replication package with the full instrument, a validated screening procedure, effect sizes reported alongside p-values, and dual-coded thematic analysis. The contribution is incremental over Gerosa et al. (2021) and Qiu et al. (2019) rather than transformative, and all evidence is single-wave, self-reported, stated-preference data from a convenience sample; the findings are correlational and exploratory in character. The paper's value depends heavily on the credibility of the reported significance structure, which currently is in doubt.","major_comments":[{"comment":"The paper states (§4.1, repeated §4.2/§4.3) that 'All findings presented in this paper were considered statistically significant, with Bonferroni adjusted p-values less than 0.05.' This is contradicted by the results tables: Table 7 includes Incentives×SE-Experience p=0.054; Table 10 includes Learning×Multilingual-Documentation p=0.051; Table 18 includes Region×Project-Age p=0.242 — all with effect sizes, inside tables of purportedly significant results. Either these rows should not be presented as findings, or the blanket claim is false. Relatedly, §4.3.2 reports a post-hoc Africa-vs-Europe difference on project age (p=0.049) while the corresponding omnibus test in Table 18 is p=0.242; pairwise post-hoc tests after a non-significant omnibus are not valid, yet this specific contrast is elevated into the Introduction as a headline result.","section":"§4.1–§4.3, Tables 4, 7, 10, 18"},{"comment":"It is unclear what 'Bonferroni adjusted' means here. If correction is only within each omnibus test's pairwise family, the RQ1 family alone is ~7 motivations × 7 demographics × 2 groups ≈ 98 omnibus tests (plus ~112 in RQ2 and more in RQ3), so marginal cells at p≈0.03–0.05 (e.g., Table 4: Learning×Gender 0.035, Learning×SE-experience 0.049) are at the expected false-positive rate. The authors must define the correction family, state whether tabled p-values are raw omnibus or adjusted post-hoc values, and either apply a family-level policy (e.g., FDR) or reframe all near-threshold results as exploratory. Text/table values also disagree: §4.1.1 text gives p=0.038 for Learning×Gender where Table 4 gives 0.035.","section":"§3.4, §4.1, Tables 4–18"},{"comment":"Several reported significant post-hoc contrasts are implausible at this sample size: §4.2.2 reports Enjoyment×Multilingual-Documentation '1% more … (p=0.027)', Networking×Web-page '3% more … (p=0.042)', Incentives×Multilingual '8% more … (p=0.024)'. With n=123 split across five intensity levels, a 1–5% difference corresponds to roughly one respondent; such a contrast cannot yield a Bonferroni-adjusted p<0.05. This strongly suggests the reported p-values are uncorrected or mis-computed. These cells should be audited and re-reported, and trivially small percentage differences should not be narrated as meaningful findings even where a test is significant.","section":"§4.2.2, Tables 12–14"},{"comment":"Headline subgroup claims rest on extremely sparse cells. Table 1: Gender 'Other' n=5, Oceania n=12, Africa n=22 (before the newcomer/practitioner split, so smaller in each analysis). Yet §4.1.2 draws Career conclusions from Female-vs-Other and Male-vs-Other contrasts, and regional claims (Africa vs America/Europe for Incentives, Reputation, Networking; RQ3 follower/project-age preferences) rest on cells of ~10–20. Kruskal–Wallis with such cells is unstable and the percentage differences (e.g., '56% more from Africa') are driven by single-digit counts. Report the exact n for every cell in Tables 3–17, and either merge, downweight, or drop contrasts involving 'Other' and Oceania.","section":"§3.3, §4.1.2, §4.3, Table 1"},{"comment":"§3.4 states Cohen's κ was computed 'following mediation' between coders. Agreement measured after discrepancies have been resolved by discussion is not inter-rater reliability — it is agreement with oneself after consensus. Additionally, the second and third authors each coded a different half of the data, so no shared subset was independently coded by all pairs, making the two κ values (0.96, 0.83) non-comparable. Either recompute κ on independent pre-mediation coding of a common subset, or remove the κ values and describe the process as negotiated consensus. This is the sole quantitative reliability evidence for RQ4.","section":"§3.4 (Thematic analysis), §4.4"}],"minor_comments":[{"comment":"Broken cross-reference: 'Table?? shows the key themes…' — the intended Table 19 is not linked.","section":"§4.4"},{"comment":"Figure 3 caption reads 'Likelihood Ratio Analysis p-values Heat Maps', but the method used throughout is Kruskal–Wallis; reconcile the caption with the text.","section":"Figure 3"},{"comment":"The abstract claims preferences for 'project age, development stage, and documentation quality vary based on specific motivations', but §4.2 reports no significant motivation differences for project age or stage among OSS practitioners. Adjust the abstract to match the RQ2 results.","section":"Abstract vs §4.2"},{"comment":"Several tables (e.g., Table 3 'Other' gender row showing 0%/0%) have empty or near-empty cells. State cell sizes in all distribution tables so percentages are interpretable.","section":"Tables 3–17"},{"comment":"Define the ε² thresholds used for 'small/medium/large' and cite the source; the labels (e.g., 0.049 'small', 0.076 'medium') imply a specific convention that is never stated.","section":"§4.1.1, Tables 4, 7"},{"comment":"Typos and consistency: 'McKight and Najab' (McKnight), 'Slighly', inconsistent 'P' vs 'p', and spacing artifacts around 'newcomers to OSS' throughout.","section":"Throughout"},{"comment":"Round 1 of Prolific recruitment had a 22% pass rate (33/150). A sentence on what this selection implies for the sample (beyond the general external-validity paragraph) would strengthen §6, e.g., whether screener-failures differ demographically from passers.","section":"§3.3, §6"},{"comment":"A substantial share of the citations are arXiv preprints (e.g., Alebachew 2025, Sesari 2025, Song 2024, Cihan 2024). Where peer-reviewed versions exist they should be cited; otherwise note the status.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The survey design, screening, and thematic procedure are competent and the topic fits the venue, but the inferential layer is currently unreliable: the blanket Bonferroni claim is refuted by the paper's own tables, several reported significant contrasts are arithmetically implausible at n=123, and at least one headline result (Africa-vs-Europe project age) is a post-hoc contrast following a non-significant omnibus test. These are fixable — the data and replication package presumably exist — but they require a full re-audit of the statistics rather than textual edits, hence major revision. I would also ask the authors to tone down the causal-sounding framing in the Discussion given the exploratory, single-wave, convenience-sample design."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: this is a clean, readable 208-person survey that actually separates newcomers (n=85) from experienced OSS people (n=123) and maps demographics → motives → stated project prefs, plus a short thematic take on recommenders. That joint comparative cut is the real increment over Gerosa, Qiu, and the usual motivation catalogues.\n\nWhat they did well is ordinary but careful. Ethics, pilot, Danilova-style programming screens, mandatory items, ε² effect sizes, dual-coded themes with high κ, Zenodo package. They keep claims correlational and give maintainers and recommender builders concrete cues (clear guidelines, source comments, small communities for some regions, interactive filters). The RQ4 themes are unsurprising but usable.\n\nThe soft spot that matters is statistics, not sampling. They say every reported finding is Bonferroni-adjusted p<0.05, yet the tables include p=0.051, 0.054, 0.242, and several ~0.03–0.05 cells that cannot survive a family of ~50–100 Kruskal–Wallis tests. A few of the headline demographic links (gender, some region cells) sit right in the noise band you expect from that many tests. RQ2 also bins people by motivation intensity without handling within-person response style, so some of the monotone Likert patterns may be yea-saying rather than motive-specific preference. Convenience sample and tiny cells (Other gender=5, Oceania=12) are real but secondary; the paper already flags external validity.\n\nNone of that makes the work incoherent. The correlation map is still directionally informative if you treat p-values as exploratory and lean on the larger effects and the newcomer/veteran split. It is for people building OSS onboarding tools or studying contributor retention—not for theory of motivation.\n\nI would send it to referees. Ask them to enforce honest multiple-comparison framing, drop or demote the non-surviving cells, and keep design advice clearly labeled as design advice. Worth engaging; not a must-cite unless you work the recommender angle.","headline":"Useful comparative survey of newcomers vs veterans on motives and project prefs, undercut by an overstated multiple-testing story.","tokens_in":27470,"tokens_out":522,"would_cite":false,"duration_ms":22898,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Who you are and why you contribute shape which open-source projects you prefer—and newcomers and veterans diverge.","keywords":["Motivations","Open Source Software","Recommendation Systems","Demographics","Newcomers","Contributor onboarding","Project selection","Survey"],"falsifier":"Track real join decisions and retention for newcomers and veterans whose stated motivations and demographics match the survey cells, and check whether they disproportionately enter projects with the preferred traits (new/active, clear guidelines, source comments, small communities, etc.) versus unmatched projects.","tokens_in":27256,"feed_emoji":"🧩","tokens_out":812,"duration_ms":27827,"temperature":0.7,"pith_summary":"Software practitioners often pick the wrong open-source projects, stall onboarding, and drop out. This survey of 208 people with software backgrounds shows that age, gender, region, education, experience, and OSS role correlate with why people contribute, and those motivations in turn correlate with preferences for project age, stage, documentation, community size, and related traits. The paper splits newcomers who have never contributed from experienced practitioners and reports distinct patterns for each group. Practitioners also say recommendation tools should be interactive, personalized to growth goals, and sensitive to collaboration style. If these links hold, maintainers and tool builders can target onboarding and retention instead of treating all contributors alike.","feed_headline":"Who you are shapes which OSS projects you pick","feed_subtitle":"Survey of 208 people finds newcomers and veterans want different project traits and better recommenders","key_machinery":"A comparative online survey of 208 screened practitioners (85 newcomers, 123 experienced), analyzed with Kruskal-Wallis tests (Bonferroni-corrected) on Likert and categorical responses linking demographics to seven motivations and eight project characteristics, plus inductive thematic analysis of open-ended recommender feedback.","core_discovery":"Demographic factors significantly correlate with OSS contribution motivations, motivations significantly correlate with stated preferences for project characteristics such as age, development stage, and documentation quality, and these patterns differ when newcomers to OSS and experienced practitioners are analyzed separately; respondents also want recommendation systems that reflect motivations, growth, and collaboration preferences.","pith_inferences":["If region and experience effects on incentives, reputation, and networking replicate, global projects may need locale-specific contribution pathways rather than English-centric defaults alone.","Cold-start recommenders that ask a short motivation questionnaire could outperform history-only models for true newcomers who have no starring or forking trail.","Merging sparse demographic cells for power may hide intersectional patterns (e.g., young women in Africa) that matter most for inclusion interventions.","Practitioner demand for LinkedIn-style cross-platform signals implies privacy and consent design will become as central as ranking quality."],"forward_implications":["Maintainers can tag tasks and repos by motivation cues (learning, helping, career) and demographic-aligned traits so the right people find them.","Recommendation systems should let users filter or converse about skills, growth goals, and collaboration style rather than only popularity or language.","Newcomer onboarding and veteran retention need different project signals; one-size defaults will miss both groups.","Recognition badges and maintainer dashboards can surface motivation-aligned opportunities without treating paid and volunteer contributors identically.","Longitudinal and domain-specific follow-ups can test whether motivation–preference links shift as people move from casual to core roles."],"fun_headline_variants":["Demographics shape OSS motivations and project picks","Newcomers and veterans want different OSS project traits","Age gender and role tie to why people join OSS projects","Motivations drive prefs for OSS age stage and docs","Survey splits new vs experienced OSS project choices"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"One-time closed-ended self-reports of motivations and hypothetical project preferences from a convenience sample after heavy screening attrition truly reflect how people choose projects in the wild.","fun_headline_variants_meta":{"raw":{"variants":["Demographics shape OSS motivations and project picks","Newcomers and veterans want different OSS project traits","Age gender and role tie to why people join OSS projects","Motivations drive prefs for OSS age stage and docs","Survey splits new vs experienced OSS project choices"]},"model":"grok-4.5","effort":"low","cost_usd":0.003446,"raw_usage":{"total_tokens":1110,"prompt_tokens":749,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":34464000,"prompt_tokens_details":{"text_tokens":749,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":306,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":749,"tokens_out":55,"duration_ms":7306,"temperature":1.0,"reasoning_tokens":306,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T00:34:17.984888+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Track real join decisions and retention for newcomers and veterans whose stated motivations and demographics match the survey cells, and check whether they disproportionately enter projects with the preferred traits (new/active, clear guidelines, source comments, small communities, etc.) versus unmatched projects.","supporting_citations":[],"review_version":1}