{"id":"54bfb274-a021-4a1f-b8af-5b0dd1d07af4","arxiv_id":"2606.03190","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Independent Chinese AI agent developers focus on user safety risks while overlooking security vulnerabilities, using informal ad-hoc practices due to lack of formal training and tools.","lead":"The interview study with 28 independent Chinese AI agent developers found they prioritize user-facing safety risks such as harmful content but show low awareness of security vulnerabilities, relying on ad-hoc safeguards and informal communication. A smart generalist might read it to see how gaps in training and platform guidance affect security practices in the growing independent AI development space.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Self-reported interview data on practices lacks validation against artifacts or observed behavior","rationale":"The reader's weakest_assumption directly names the methodological hinge on which the descriptive claims about practices rest. No stronger internal inconsistency or hidden assumption appears in the abstract or stated claims; the paper is explicitly positioned as an exploratory interview study rather than a validated measurement of behavior.","tokens_in":1739,"tokens_out":309,"duration_ms":14027,"concrete_test":"Identify 5–10 public GitHub repositories of AI agent projects by independent Chinese developers (via language, README, or author signals); manually inspect for presence/absence of formal S&P mechanisms (e.g., input sanitization libraries, auth frameworks, logging of security events) versus purely ad-hoc prompt engineering or manual filters. If the distribution of formal vs. ad-hoc implementations deviates substantially from the interview themes, the self-report accuracy claim is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that developers exhibit low awareness of security vulnerabilities and rely exclusively on ad-hoc manual safeguards—depends entirely on participants' verbal descriptions of their own understandings and actions. In an interview-only design with no code review, log analysis, or behavioral observation, recall bias, social desirability, or incomplete probing could systematically distort reports of awareness and practice. The N=28 convenience sample of Chinese developers who agreed to discuss S&P further compounds this, as self-selection may favor those with stronger (or weaker) practices than the population average.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper reports results from a qualitative interview study with N=28 independent Chinese AI agent developers. It claims that these developers adopt a user-centric perspective, prioritizing user-facing safety risks such as harmful content while showing low awareness of security vulnerabilities; as a result they rely exclusively on ad-hoc, manually crafted safeguards and informal communication rather than formal tools or processes. The study identifies primary inhibitors as lack of formal S&P training, accessible security tools, and actionable platform guidance, positioning the work as the first exploration of this population's S&P understandings.","tokens_in":1827,"tokens_out":546,"duration_ms":20503,"significance":"If the findings are reliable, the work supplies the first empirical account of S&P practices among independent AI agent developers, a growing population outside corporate structures. The identification of a user-safety versus technical-security gap and the listed inhibitors could directly inform the design of targeted tooling and educational resources for this group.","major_comments":[{"comment":"Methods section: The manuscript supplies no details on recruitment procedures, interview protocol, transcription/coding process, or inter-coder reliability. Without these, it is impossible to evaluate selection bias in the convenience sample or the rigor with which low security awareness versus high user-safety focus was probed, directly undermining confidence that the data support the central claims.","section":"Methods"},{"comment":"Findings / Results sections: All claims of “low awareness of security vulnerabilities” and “absence of formal tools or processes” rest exclusively on self-reported interview responses. No triangulation via code inspection, artifact analysis, or behavioral observation is described; this is load-bearing because recall bias or social-desirability effects could systematically inflate reports of ad-hoc practices and understate actual tool use.","section":"Findings"},{"comment":"Discussion: The assertion that the N=28 sample suffices to identify the “main challenges and inhibitors for the broader population” is not supported by any discussion of saturation, sample diversity, or limitations of self-selection among developers willing to discuss S&P topics.","section":"Discussion"}],"minor_comments":[{"comment":"The abstract and introduction repeatedly use “S&P” without an initial definition; a brief parenthetical expansion on first use would improve readability.","section":"Abstract"},{"comment":"Related-work section appears to omit several recent HCI studies on developer security practices outside the AI-agent context; adding 2–3 citations would better situate the novelty claim.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback. We address each major comment below and indicate where revisions will be incorporated.","responses":[{"response":"We agree the methods section is insufficiently detailed. In revision we will expand it with: recruitment via WeChat groups, GitHub issues and snowball sampling among independent Chinese AI developers; the full semi-structured interview protocol with example prompts on user-safety vs. technical security; transcription process; inductive thematic coding procedure; and inter-coder reliability steps (two coders, discrepancy resolution meetings). These additions will allow readers to assess selection bias and probing rigor.","revision_made":"yes","referee_comment":"[Methods] Methods section: The manuscript supplies no details on recruitment procedures, interview protocol, transcription/coding process, or inter-coder reliability. Without these, it is impossible to evaluate selection bias in the convenience sample or the rigor with which low security awareness versus high user-safety focus was probed, directly undermining confidence that the data support the central claims."},{"response":"The study is an exploratory qualitative interview investigation of reported understandings and practices; self-report is the primary data source by design. We will add an explicit limitations paragraph discussing recall bias and social-desirability effects. Because the original protocol did not include artifact collection, we cannot retroactively add triangulation; we will note this as a boundary condition and suggest future mixed-methods work.","revision_made":"partial","referee_comment":"[Findings] Findings / Results sections: All claims of “low awareness of security vulnerabilities” and “absence of formal tools or processes” rest exclusively on self-reported interview responses. No triangulation via code inspection, artifact analysis, or behavioral observation is described; this is load-bearing because recall bias or social-desirability effects could systematically inflate reports of ad-hoc practices and understate actual tool use."},{"response":"We will revise the discussion to report thematic saturation (no new codes after the 22nd interview), sample diversity (range of developer experience, project types, and locations), and self-selection limitations. This framing is consistent with qualitative standards for identifying core themes in an emerging population rather than claiming statistical generalizability.","revision_made":"yes","referee_comment":"[Discussion] Discussion: The assertion that the N=28 sample suffices to identify the “main challenges and inhibitors for the broader population” is not supported by any discussion of saturation, sample diversity, or limitations of self-selection among developers willing to discuss S&P topics."}],"tokens_in":1408,"tokens_out":540,"duration_ms":10463,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that independent Chinese AI agent developers prioritize user-facing safety like harmful content but have low awareness of security vulnerabilities, leading them to use only ad-hoc manual safeguards.\n\nThis is new because it's the first study focused on this population of independent devs in China relying on global LLM services. Prior work has looked at corporate settings or other regions, but not this combination.\n\nThe paper does a solid job pulling out the specific challenges from the interview data, like lack of formal training and accessible tools, and showing how developers think from the user's perspective.\n\nThe soft spot is the interview-only design. Claims about actual practices rest on what people said in interviews, with no code review or behavioral checks to confirm. The N=28 convenience sample also means self-selection could skew the results toward certain types of developers.\n\nThis paper is for HCI and usable security researchers working on AI developer support. Readers looking for insights into emerging developer populations will find the reported practices and gaps useful.\n\nIt deserves serious peer review because the topic is relevant and the qualitative approach surfaces actionable points, even with the self-report limits. I'd recommend sending it out rather than rejecting at the desk.","headline":"Independent Chinese AI devs focus on user safety but miss security risks, based on interviews that could use more validation.","tokens_in":2304,"tokens_out":307,"would_cite":false,"duration_ms":17419,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Independent Chinese AI agent developers focus on user-facing safety risks like harmful content but show low awareness of security vulnerabilities.","keywords":["AI agents","security and privacy","independent developers","user perspective","ad-hoc safeguards","interview study","China","safety risks"],"falsifier":"A survey or audit of a larger group of independent developers finding that most use formal security tools and processes in their AI agents.","tokens_in":2641,"feed_emoji":"🔒","tokens_out":560,"duration_ms":23534,"temperature":0.7,"pith_summary":"An interview study with 28 Chinese independent developers reveals they approach security and privacy primarily from the user's viewpoint. They emphasize protecting against harmful outputs while paying little attention to underlying vulnerabilities in their systems. As a result, they depend on custom manual safeguards and informal talks rather than established tools or procedures. These habits arise mainly from insufficient training, scarce security resources, and vague platform advice. Understanding this helps explain potential weaknesses in the growing number of AI agents built outside big companies.","feed_headline":"AI devs focus on user risks but miss security threats","feed_subtitle":"Interviews show Chinese independent creators rely on manual fixes due to missing training and tools.","key_machinery":"The user-centric mindset directing attention to content safety over technical security vulnerabilities in AI agent development.","core_discovery":"Independent developers frequently think and act from their users' perspective. They focused on user-facing safety risks such as harmful content while exhibiting low awareness of security vulnerabilities. Consequently, developers rely almost exclusively on ad-hoc, manually crafted safeguards and informal communication, with an absence of formal tools or processes for S&P practices. These actions are driven by a lack of formal training on S&P related skills, accessible security tools and actionable guidance from platforms.","pith_inferences":["Security researchers might find more vulnerabilities in independently developed AI agents than in corporate ones.","Efforts to improve AI safety could benefit from targeting individual developers with educational resources.","Similar patterns may exist among independent developers in other regions using global LLM services."],"forward_implications":["Developers create AI agents with ad-hoc protections that may miss common security issues.","Absence of formal S&P processes increases reliance on personal judgment.","Platforms could address gaps by offering better guidance and tools.","Lack of training leads to uneven S&P practices across independent projects."],"fun_headline_variants":["Chinese AI devs focus on user safety overlook security","Independent devs rely on manual security fixes","No formal tools for AI agent security practices","Lack of training on security for Chinese AI devs"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That the self-reported behaviors of the 28 interviewed developers accurately mirror their actual practices and represent the wider population of independent Chinese AI agent developers.","fun_headline_variants_meta":{"raw":{"variants":["Chinese AI devs focus on user safety overlook security","Independent devs rely on manual security fixes","No formal tools for AI agent security practices","Lack of training on security for Chinese AI devs"]},"model":"grok-4.3","cost_usd":0.01158,"raw_usage":{"total_tokens":5067,"prompt_tokens":655,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":115799500,"prompt_tokens_details":{"text_tokens":655,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4358,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":655,"tokens_out":54,"duration_ms":30365,"temperature":1.0,"reasoning_tokens":4358,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T08:46:58.555681+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A survey or audit of a larger group of independent developers finding that most use formal security tools and processes in their AI agents.","supporting_citations":[],"review_version":1}