{"id":"49d0fd9d-bbe2-4860-aa45-e7c6bb08af93","arxiv_id":"2505.04277","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A participatory design fiction study with 19 OSS newcomers yields 32 AI mentor design strategies and shows the biggest research gaps are in project discovery and structure comprehension.","lead":"Newcomers to open-source projects were asked to imagine an AI mentor that guides them from picking a project to merging their first code, and then tested a prototype of that concept. The result is 32 design strategies, with the strongest demands in project discovery and understanding codebases, areas where current research is thin.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Effectiveness claim rests on self-reported ratings of a non-functional prototype by the same 19 participants who proposed the strategies; no objective onboarding outcome was measured.","rationale":"The reader's weakest assumption identifies the same core issue: the 19 participants proposed the strategies and then rated the prototype, so the validation is partly self-confirmation. My stress-test adds two concrete refinements: (1) Section 3.2.1 explicitly describes the prototype as simulating the onboarding process rather than implementing it, so the TAM responses in Section 4.3.3 are reactions to a mockup, not to a usable system; and (2) the paper never measures an objective onboarding outcome, so phrases like 'demonstrating the relevance and effectiveness of the proposed strategies' (Section 6) go beyond what the data can support. This does not invalidate the study's exploratory contribution: the design fiction method is appropriate for eliciting needs, the 32 strategies are clearly grounded in participant quotes, and the literature review identifies plausible research gaps. The manuscript already self-identifies as a first step, and the reader's CONDITIONAL verdict appropriately requires independent evaluation with a broader population and a working system. My concern strengthens that condition but does not move the verdict to REJECT or ACCEPT, because the central catalog contribution stands on its own and the paper's own framing in Section 3 positions Phase II as a suitability study rather than a full effectiveness trial. Therefore the verdict should remain CONDITIONAL, with the condition made explicit: effectiveness claims require replication with naive users and objective measures.","tokens_in":24521,"tokens_out":1906,"duration_ms":22513,"concrete_test":"Run a between-subjects evaluation with fresh OSS newcomers who did not participate in the design fiction sessions: randomly assign one group to a functional OSSerCopilot prototype (or a high-fidelity Wizard-of-Oz version that actually supports the simulated steps) and another group to standard GitHub workflows with good-first-issue labeling. Measure objective onboarding outcomes--time to first acceptable pull request, task completion rate, appropriateness of chosen issues, and sustained contribution after 30 days--alongside the same TAM survey. If the TAM ratings do not replicate above baseline or the objective outcomes do not differ significantly, the 'effectiveness' claim is unsupported and the verdict should remain conditional on future validation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that the 32 design strategies are 'useful and effective' (Abstract, Section 6) depends on Phase II evidence, but that evidence has a structural gap. Section 3.2.1 states the prototype 'simulates the entire newcomer onboarding process'; it is a demonstration with a PyTorch example, not a working AI mentor. Section 4.3.3 then reports TAM ratings from essentially the same 19 participants who co-designed the strategies in Phase I. All items on usefulness received over 88% agreement and all participants agreed they would use OSSerCopilot in the future (S1). This is a classic self-confirmation loop: participants evaluate a mockup of their own elicited ideas, with no baseline, no independent users, no actual task performance, and no measure of whether onboarding outcomes improve. TAM measures perceived usefulness and intention, not realized effectiveness; willingness to use a simulated tool cannot establish that a real AI mentor would reduce time-to-first-contribution, improve task selection, or reduce dropout. The paper itself does not claim the prototype was functional, yet the conclusion language ('demonstrating the relevance and effectiveness of the proposed strategies') overreaches. Threats to validity in Section 5 discuss platform generalizability and literature coverage, but do not address the circularity of having the same participants validate their own design proposals, nor the absence of objective outcome data. This is the most load-bearing weakness because if the effectiveness claim is removed, the paper's primary contribution reduces to a useful catalog of newcomer expectations and research gaps, which is still valuable but not a validated AI mentor design.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper aims to explore the potential of an AI mentor to support newcomers throughout the entire OSS onboarding process. Using Design Fiction as a participatory method, the authors conducted sessions with 19 OSS newcomers, elicited 32 design strategies for an AI mentor across nine onboarding steps, and then built a prototype called OSSerCopilot that implements these strategies. They evaluated the prototype through semi-structured interviews and a Technology Acceptance Model survey, and they compared the design strategies against a literature review of 537 papers to identify research gaps. The reported findings are that newcomers face the greatest difficulties in discovering projects, grasping project structure, and identifying tasks; that the prototype's usefulness and ease of use were rated highly by participants; and that early onboarding steps such as project discovery and structure understanding are underexplored in current research.","tokens_in":24839,"tokens_out":3466,"duration_ms":36254,"significance":"If taken as a design-space exploration, this paper makes a useful and novel contribution: it is the first study, to my knowledge, to systematically envision an AI mentor spanning the entire OSS onboarding process, and the 32 design strategies are concrete, actionable, and grounded in a participatory method. The literature review in RQ4 is also valuable as a map of research gaps, and the paper is transparent about its qualitative analysis procedures, information saturation, and data availability. The authors are to be credited for sharing the fiction story, prototype, and codebook. However, the paper's central claim that the design strategies are 'useful and effective' is not supported by the evidence presented, because the validation relies on self-reported perceptions of a non-functional prototype by the same participants who proposed the strategies.","major_comments":[{"comment":"The validation of the design strategies is circular. The 19 participants who proposed the 32 strategies in Phase I are essentially the same participants who rated the prototype in Phase II (Section 3.2.2 states that all participants except P5 joined the prototype feedback sessions). These participants evaluated a prototype that implements their own elicited ideas, which means the high TAM ratings (all usefulness items over 88% agreement, and S1 at 100%) partly confirm the participants' own earlier aspirations. Because Section 3.2.1 explicitly states that the prototype 'simulates the entire newcomer onboarding process' and is demonstrated with a single PyTorch example, rather than being a functioning AI mentor, the TAM results cannot establish that the strategies would improve real onboarding outcomes such as time-to-first-contribution, task selection quality, or newcomer retention. An evaluation by independent participants, or a baseline comparison, or objective task-performance measures is needed to support the effectiveness claim.","section":"Section 3.2.1 and Section 4.3.3"},{"comment":"The wording 'demonstrating the relevance and effectiveness of the proposed strategies' (Section 6) and 'which suggests the design strategies are effective' (Abstract) overreaches the evidence. The data support only that the co-designing participants perceived the prototype as useful and easy to use. Since the prototype is a simulation and the evaluation is a perception study, the conclusion should be reframed as evidence of perceived usefulness and intention to use, not realized effectiveness. This is a load-bearing issue because the stated contribution of the paper is the effectiveness of the strategies.","section":"Abstract, Section 4.3.3, and Section 6"},{"comment":"The Threats to Validity section discusses generalizability, information saturation, and literature collection, but it does not address the most significant threat to the RQ3 conclusion: having the same participants validate their own design proposals, nor the absence of objective outcome measures. The 'Reliability of results' subsection emphasizes the TAM reliability check and coding procedures, but peripheral validation steps do not mitigate the structural circularity of the Phase II evaluation. The authors should explicitly acknowledge that the prototype evaluation is a self-confirmation check and therefore can only be interpreted as a refinement of the elicited strategies, not as a test of their real-world effectiveness.","section":"Section 5 (Threats to Validity)"},{"comment":"The sample is narrow for the generalizability claim implicit in the paper's conclusions. Of the 19 participants, 14 are undergraduate students, 17 are male, and 13 have at most two commits; most were recruited from a single university OSS course and local communities. While the paper acknowledges platform generalizability in Section 5, it does not address how the demographic and experience distribution might bias the elicited strategies toward the needs of novice students rather than the broader OSS newcomer population. This matters because the literature-gap analysis in RQ4 is only as useful as the representativeness of the strategy set; the authors should either temper the population-level conclusions or provide additional evidence of transferability.","section":"Table 1 and Section 3.1.3"}],"minor_comments":[{"comment":"The prototype name is misspelled as 'OSSerCopliot' in the paragraph beginning 'We implemented our prototype as a Web plugin'; it should be 'OSSerCopilot'.","section":"Section 4.3.1"},{"comment":"The phrase 'conducing a more comprehensive difficulty assessment' contains a typo; it should read 'conducting'.","section":"Section 4.2.2, 'Identify a Task'"},{"comment":"The radar chart in Figure 4 compares expectation scores (average ranking scores, 1-9 scale from Section 4.2.1) with satisfaction ratings (Likert scale, 1-5 scale from Section 4.3.2). Because these variables use different scales, the visual overlap may be misleading. The authors should either normalize the two scales or clarify that the figure is only a qualitative comparison.","section":"Section 4.3.2, Figure 4"},{"comment":"The text mentions that participants use 'GFI websites' to search for issues, but no citation or reference is provided for these websites; adding a citation or a footnote would improve traceability.","section":"Section 4.2.2, 'Identify a Task'"},{"comment":"The literature review methodology is based on Google Scholar searches with the top 100 papers per keyword, which is not a fully reproducible systematic mapping. Although the authors acknowledge this limitation in Section 5, reporting the exact search dates and keyword combinations in the supplemental material would improve the reproducibility of the RQ4 gap analysis.","section":"Section 3.3 (Phase III)"}],"recommendation":"major_revision","confidential_remarks":"The paper is better positioned as a participatory design study whose main deliverable is the set of 32 strategies and the research-gap map, rather than as an empirical demonstration of an effective AI mentor. The current claims in the abstract and conclusion exceed the evidence. I suggest the editor require the authors to (1) reframe the effectiveness language consistently, (2) add a dedicated discussion of the self-confirmation threat and the absence of outcome measures in the Threats to Validity section, and (3) consider whether a brief independent evaluation (even with a small number of external participants or a functional prototype with a simple task) would strengthen the RQ3 claim enough to avoid overstatement. The paper is suitable for the venue if these issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core of this paper is the 32 design strategies for an AI mentor spanning the whole OSS onboarding workflow, elicited via participatory design fiction from 19 newcomers. That is genuinely new as a comprehensive whole-process view, and it comes with a prototype, a codebook, and a literature gap analysis. The methods section is transparent, the qualitative coding is described in enough detail to be audited, and the supplementary materials are available. Credit where due: this is careful qualitative work, and the gap analysis points to real under-studied steps like project discovery and structure comprehension.\n\nThe soft spot is exactly where the stress-test puts it. The central claim that the strategies are \"effective\" rests on Phase II TAM ratings of a prototype that simulates rather than implements the AI mentor, filled in by the same participants who proposed the strategies in Phase I. All usefulness items got over 88% agreement and all participants said they would use it. That is a self-confirmation loop. There is no baseline, no independent users, no objective measure of onboarding outcomes, and no working system. TAM measures perceived usefulness and intention, not realized effectiveness. Willingness to use a mockup of your own ideas does not demonstrate that a real AI mentor would shorten time-to-first-PR or reduce dropout. The paper itself does not claim the prototype is functional, yet the conclusion language—\"demonstrating the relevance and effectiveness of the proposed strategies\"—overreaches. The threats-to-validity section covers platform generalizability and literature coverage but never mentions this circularity or the absence of outcome data. That is the load-bearing weakness.\n\nOther soft spots are minor in comparison. The sample is mostly male undergraduates from one course and local communities, though the authors acknowledge generalizability limits. The literature review used Google Scholar with top-100-per-keyword filtering, which could miss relevant work; they list this as a threat, and the keyword sets look reasonable. The title and repeated use of \"revolutionize\" is marketing, not substance, but it does not affect the underlying data.\n\nWho is this for? Researchers and tool builders working on AI support for OSS newcomers, and anyone designing onboarding interventions. The strategy catalog and gap analysis are worth having. But read it as an exploratory needs analysis with design implications, not as validated evidence of effectiveness.\n\nRecommendation: this deserves peer review as a qualitative study. A serious referee should ask for the effectiveness conclusion to be softened to perceived usefulness/suitability, and for an explicit limitation paragraph on the same-participant circularity and lack of objective validation. With that reframing, it is a legitimate FSE/ICSE-style empirical paper. I would not desk-reject it.","headline":"Solid exploratory design study with a useful catalog of 32 AI-mentor strategies for OSS onboarding, but the effectiveness claim rests on a circular self-report of a non-functional prototype by the same 19 participants who proposed the strategies.","tokens_in":25305,"tokens_out":1297,"would_cite":true,"duration_ms":15813,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an AI mentor can cover the entire OSS newcomer onboarding process, that newcomers co-designed 32 strategies for it, and that a prototype implementing those strategies was rated useful and easy to use.","keywords":["open source software","newcomer onboarding","AI mentor","design fiction","participatory design","technology acceptance model","good first issues","GitHub"],"falsifier":"Deploy a working OSSerCopilot (or its top-priority strategies) with a random set of real newcomers and compare first-contribution completion and retention against a control group using conventional onboarding resources; no improvement in actual contribution behaviour would refute the claim that the strategies are effective. A broader, more representative survey that finds the 32 strategies miss key newcomer needs would also weaken it.","tokens_in":1685,"feed_emoji":"🤖","tokens_out":2062,"duration_ms":81997,"temperature":0.7,"pith_summary":"The paper argues that the full journey of an open-source newcomer, from picking a project to getting a pull request merged, can be mentored by artificial intelligence, and that newcomers themselves can specify how such mentoring should work. Through participatory Design Fiction, 19 OSS newcomers produced 32 design strategies for an \"AI mentor\" covering every step of the contribution workflow. The authors built OSSerCopilot, a GitHub-integrated prototype implementing the strategies, and found high perceived usefulness, ease of use, and stated willingness to use it. The claim matters because human expert mentoring does not scale, and onboarding dropout is a known threat to open-source sustainability. A literature review adds that current AI research clusters on coding and testing while neglecting the early steps newcomers most want help with, so the strategies point to open research opportunities.","feed_headline":"Newcomers want an AI mentor at every step of open-source onboarding","feed_subtitle":"Nineteen OSS newcomers shaped 32 design strategies; every one said they would use the prototype.","key_machinery":"The machinery is the AI mentor design-strategy set, produced by Design Fiction, a participatory method in which participants watch a short animated fiction set in 2030 and then discuss, at each of nine contribution steps, how an ideal AI should help them. The 32 strategies are the specification, and OSSerCopilot is their implementation: a web plugin integrated into GitHub with a progress bar, a customized project-recommendation form, project-suitability analysis, issue-difficulty sorting, pull-request submission simulation, and feedback summarization. The Technology Acceptance Model, a standard survey model measuring perceived usefulness, ease of use, and self-predicted future use, is the measuring instrument that carries the validation.","core_discovery":"On the paper's own terms, the discovery is that a whole-process AI mentor for OSS onboarding is a designable, acceptable, and currently under-supported idea. The 19 participants described their current onboarding as self-service, relying on search engines and community resources rather than direct community help, and flagged \"grasp the project's structure,\" \"discover an interested project,\" and \"identify a task\" as the hardest steps. They then articulated 32 design strategies, from personalized project recommendation and issue-difficulty assessment to pull-request submission simulation and reviewer-feedback summarization. The prototype OSSerCopilot put those strategies into a GitHub-sidebar plugin, and participants' Technology Acceptance Model responses showed strong agreement on usefulness and ease of use, with all participants saying they would use it; the authors take this as evidence the strategies are effective. The accompanying literature review found only one paper each for project discovery and contribution-guideline support, against 231 for coding and 129 for testing, identifying the early steps as the largest gap between newcomer expectations and existing research.","pith_inferences":["The strongest test the paper leaves undone is deployment: give real newcomers a working OSSerCopilot and compare actual first-contribution rates and retention against a control group, since self-reported acceptance can overstate real use.","The strategy set implies a research rebalancing: effort on project discovery and repository-level comprehension may produce larger onboarding gains than further code-generation improvements, given how little existing work targets those steps.","The issue-difficulty strategy has an isolable prediction: an AI that ranks issues by difficulty should beat current \"good first issue\" labels at selecting tasks newcomers actually complete, which could be tested on historical issue data.","Participants' preference for GitHub integration suggests platform-embedded mentoring will be adopted more readily than standalone tools, but the minority who wanted access to local code files points to a hybrid deployment worth exploring."],"forward_implications":["Tool builders can take the 32 strategies as a direct interface specification for an AI mentor, since each strategy maps onto concrete features such as recommendation forms, issue-difficulty ranking, and PR simulation.","The highest-value targets are the early steps, discovering a project, grasping its structure, and identifying a task, where newcomer demand is strongest and existing research is thinnest.","Newcomers want an integrated assistant inside GitHub rather than a separate tool, and they want it to explain and guide rather than write code for them, preserving their learning.","OSS maintainers should treat \"good first issue\" labels as only a starting point, because AI-assisted difficulty assessment could expand the set of suitable tasks beyond labeled issues.","If the strategies are built into practice, expert mentoring workload should drop because the AI absorbs the repetitive parts of onboarding, leaving experts for higher-level review."],"supporting_citations":[{"why":"Supplies the nine-step contribution workflow that structures the fiction story and the design strategies.","marker":"[17]"},{"why":"Defines Design Fiction, the participatory method used to elicit newcomer needs.","marker":"[10]"},{"why":"Demonstrates Design Fiction for tool design in OSS pull requests; its session format is adapted here.","marker":"[73]"},{"why":"Provides the Technology Acceptance Model used to measure perceived usefulness, ease of use, and future use.","marker":"[18]"},{"why":"Systematically catalogues newcomer barriers, motivating the need for an AI mentor.","marker":"[60]"},{"why":"Shows good first issues are often unsuitable for newcomers, motivating the issue-difficulty assessment strategies.","marker":"[68]"},{"why":"Analyses mentoring on good first issues and identifies newcomer challenges the AI mentor addresses.","marker":"[66]"},{"why":"Earlier proposal of a chatbot mentor for newcomers that this work extends to a whole-process AI mentor.","marker":"[21]"}],"fun_headline_variants":["All 19 OSS newcomers say yes to AI mentor for every onboarding step","OSSerCopilot: AI mentor for OSS onboarding gets unanimous yes","AI mentor for OSS onboarding: 19/19 users on board, research lags","OSS onboarding: AI mentor design accepted, but early steps lack research"],"cache_read_input_tokens":27392,"weakest_assumption_plain":"The load-bearing premise is that 19 participants, mostly male undergraduates from one course and nearby communities, speak for the wider population of OSS newcomers, and that their positive ratings of a prototype they helped design predict real onboarding success.","fun_headline_variants_meta":{"raw":{"variants":["All 19 OSS newcomers say yes to AI mentor for every onboarding step","OSSerCopilot: AI mentor for OSS onboarding gets unanimous yes","AI mentor for OSS onboarding: 19/19 users on board, research lags","OSS onboarding: AI mentor design accepted, but early steps lack research"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001363,"raw_usage":{"total_tokens":5583,"prompt_tokens":1056,"completion_tokens":4527,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":4444}},"tokens_in":672,"tokens_out":4527,"duration_ms":29985,"temperature":1.0,"reasoning_tokens":4444,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:31:39.017116+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deploy a working OSSerCopilot (or its top-priority strategies) with a random set of real newcomers and compare first-contribution completion and retention against a control group using conventional onboarding resources; no improvement in actual contribution behaviour would refute the claim that the strategies are effective. A broader, more representative survey that finds the 32 strategies miss key newcomer needs would also weaken it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates Design Fiction for tool design in OSS pull requests; its session format is adapted here."}],"review_version":1}