{"id":"ae27554b-e1c0-4952-a062-f6cc9cc21d9b","arxiv_id":"2607.11309","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Bundled PHITS knowledge bases and agent policies let general AI tools edit inputs, run simulations, optimize parameters, compile code, and interpret results across five demo workflows.","lead":"The authors package PHITS manuals, samples, and rules so general AI assistants and agents can prepare inputs, run Monte Carlo particle-transport jobs, fix errors, and interpret results. This shows a practical path to lower the expertise barrier for radiation-transport simulation without building a custom AI product.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The five author-curated demos do not establish that the bundled resources enable reliable general AI-agent PHITS workflows for typical users without frequent human oversight.","rationale":"The reader’s weakest_assumption correctly isolates the load-bearing gap: the five demos, performed under author-curated resources and still requiring human verification, are not shown to be representative of real PHITS usage. The abstract alone offers only qualitative success statements, so the CONDITIONAL verdict with low confidence remains appropriate. No internal inconsistency or circularity is present; the paper is an infrastructure/demo contribution whose soundness hinges on evaluation depth the abstract does not demonstrate. Shipping the resources together with quantitative and held-out metrics would move the work toward ACCEPT; their absence keeps the verdict unchanged.","tokens_in":2066,"tokens_out":459,"duration_ms":15420,"concrete_test":"In the full paper, verify whether the evaluation section reports (i) quantitative success/failure or human-intervention rates over repeated independent runs of each of the five tasks and (ii) results on at least one held-out workflow never used to curate the knowledge base or policies; if either is absent or interventions remain frequent, the claim that the resources enable reliable agent-driven workflows does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that packaging a RAG knowledge base plus a compact agent reference with PHITS-specific policies and execution rules lets general-purpose assistants (NotebookLM, Codex, Claude Code) edit inputs, run calculations, inspect errors, analyze results, and assist with source modification across complex multi-step workflows—rests entirely on five demonstration tasks. The abstract reports that the agents “could handle” those tasks when the curated resources were supplied, yet simultaneously lists “human verification” as a practical lesson and supplies no success rates, intervention counts, failure-mode analysis, no-resource baselines, or evidence that the five tasks span the distribution of real PHITS user work. Under those conditions the demos remain existence proofs under author-controlled settings rather than support for the broader claim of reliable, general agent-driven particle-transport practice.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes a code-side strategy for applying general-purpose AI assistants and agents to the Monte Carlo particle-transport code PHITS, rather than building a dedicated PHITS-specific AI application. The authors prepare two complementary AI-ready resource sets from manuals, lecture materials, sample inputs, utilities, and developer cautions: (i) a bundled knowledge base for RAG-based assistants (demonstrated with NotebookLM) and (ii) a compact agent reference combined with PHITS-specific policies and execution rules (demonstrated with Codex and Claude Code). Across five demonstration tasks—input modification, repeated simulations, parameter optimization, source modification/compilation, post-processing, and result interpretation—the abstract reports that agents could edit inputs, execute calculations, inspect errors, analyze results, and assist with compilation when those resources and rules were supplied. Practical lessons listed include precise prompts, human verification, well-documented samples, explicit execution policies, and CLI-accessible tools. The central claim is that bundling such resources with the code enables general-purpose AI tools for complex PHITS workflows.","tokens_in":2232,"tokens_out":1327,"duration_ms":14471,"significance":"If substantiated, the work would be a useful engineering contribution for the computational nuclear/particle-transport community: it reframes AI integration as a packaging and documentation problem (RAG knowledge base + agent reference + policies) rather than a custom-model problem, and it is immediately actionable for PHITS users and maintainers. Credit is due for the concrete resource design (dual RAG vs. agent-reference tracks), the multi-tool demonstration scope (NotebookLM, Codex, Claude Code), and the explicit inclusion of compilation and source-modification workflows, which go beyond simple chat Q&A. The practical lessons (execution policies, CLI tools, sample quality) are transferable to other Monte Carlo codes. Significance, however, hinges on whether the five demos are shown to be representative and reproducible with quantified reliability, not only existence proofs under author-curated conditions.","major_comments":[{"comment":"The abstract’s central claim—that AI agents “could handle complex PHITS workflows when appropriate resources and rules were provided”—rests on five demonstration tasks with no reported success rates, intervention counts, failure modes, or error bars. For a computational-physics methods paper this is load-bearing: without quantitative outcomes the demos remain existence proofs under author-controlled settings and do not yet support the broader claim of reliable agent-driven practice. Please add per-task metrics (success/partial/fail, number of human interventions, wall-time, token/cost if relevant) and a short failure-mode taxonomy.","section":"Abstract (Results / five demonstration tasks)"},{"comment":"No no-resource or weak-resource baseline is described. The claim that bundling the RAG knowledge base and agent reference is what enables the workflows cannot be isolated from general LLM capability unless the same five tasks are also attempted without those resources (or with only the public manual). A controlled ablation—or at least a qualitative side-by-side—is needed to make the packaging recommendation evidence-based rather than anecdotal.","section":"Abstract (resource design and results)"},{"comment":"“Human verification” is listed among the practical lessons while the abstract simultaneously asserts that agents handled complex multi-step workflows. This tension is load-bearing for the reliability claim. Please clarify the intended autonomy level (fully autonomous vs. human-in-the-loop), state where verification was required in each demo, and revise the claim language so that it matches the actual degree of oversight used.","section":"Abstract (practical lessons)"},{"comment":"Representativeness of the five tasks is not established. The weakest assumption of the paper is that success on these author-chosen demos generalizes to typical PHITS user work. Please justify task selection against common PHITS use cases (e.g., shielding, activation, medical, accelerator), state what was deliberately excluded (geometry complexity, variance reduction, multi-code coupling, large tallies), and discuss limits of generalization.","section":"Abstract (five demonstration tasks)"}],"minor_comments":[{"comment":"The abstract packs many workflow stages into one sentence (“input modification, repeated simulations, parameter optimization, program compilation, post-processing, and result interpretation”). A short table or enumerated list mapping each demo to tools used (NotebookLM vs. Codex vs. Claude Code) and resources consumed would improve clarity for readers scanning the abstract.","section":"Abstract"},{"comment":"Terms such as “AI-ready resources,” “compact agent reference,” and “PHITS-specific policies and execution rules” are central but undefined in the abstract. Brief parenthetical definitions (what files, what format, approximate size) would help non-AI specialists in the PHITS community.","section":"Abstract"},{"comment":"The abstract asserts support for “bundling AI-ready resources with particle transport codes” in general, while all evidence is PHITS-specific. Soften the cross-code generalization in the closing sentence or flag it as a hypothesis pending similar packaging for other codes (e.g., MCNP, Geant4, OpenMC).","section":"Abstract (closing sentence)"},{"comment":"If the full manuscript includes the knowledge-base construction pipeline, sample prompts, and execution-policy text, please ensure they are archived (supplement or repository) so that other groups can reproduce the agent setup; the abstract alone does not indicate availability.","section":"Abstract / data availability (expected in full text)"}],"recommendation":"major_revision","confidential_remarks":"Review is based solely on the abstract (full text not available). Confidence in the recommendation is therefore limited: if the full paper already contains quantitative success rates, baselines, and failure analysis, several major comments may reduce to minor. Scope fit for physics.comp-ph is reasonable as an engineering/methods note, but the journal should expect the quantitative evaluation bar of a methods paper rather than a pure demo write-up. No concerns about misconduct or citation pattern from the material provided."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a code-side packaging paper, not a transport-physics result. The authors curate two PHITS-ready resource sets (RAG knowledge base plus a compact agent reference with policies/execution rules) so off-the-shelf tools—NotebookLM, Codex, Claude Code—can edit inputs, run jobs, chase errors, post-process, and even help with compile/source work. That packaging pattern is the real contribution.\n\nWhat is new and useful is the deliberate dual-resource design and the PHITS-specific policies, not the idea of RAG or coding agents in the abstract. Shipping manuals, samples, cautions, and execution rules as first-class AI-facing artifacts is a practical move other Monte Carlo codes could copy. The five demos span a sensible range (input edits, repeated runs, optimization, compilation, post-processing/interpretation). As an existence proof under author-curated resources, that is fair engineering work. Circularity is low: they are not fitting free parameters to recover a law; they are demonstrating a workflow stack.\n\nSoft spots, in proportion: the abstract only says agents “could handle” the tasks and immediately lists human verification as a lesson. No success rates, intervention counts, failure modes, no-resource baselines, or evidence that the five tasks represent typical PHITS user work. The stress-test note is right that this does not yet establish reliable general agent-driven practice—only that the stack can work when the authors set the table. That is a real but expected limitation for an abstract-only methods note, not a load-bearing math flaw.\n\nWho it is for: PHITS maintainers, Monte Carlo tool builders, and people building AI-assisted scientific software who want a concrete packaging recipe. Not for someone hunting new transport physics. It deserves a serious referee if the full paper ships the resources (or a clear description of them) and an honest evaluation section; desk-rejecting pure infrastructure demos of this kind would be a mistake. I would send it to review as applied computational methods, expecting pressure for metrics and failure analysis rather than rejection on novelty of the underlying AI tools.","headline":"Solid engineering packaging for PHITS+AI agents; existence demos, not yet a reliability claim.","tokens_in":2871,"tokens_out":521,"would_cite":false,"duration_ms":8203,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"PHITS can be driven by general AI agents once shipped with two curated resource packs and execution rules.","keywords":["PHITS","Monte Carlo particle transport","AI agents","retrieval-augmented generation","AI-ready resources","workflow automation","code-side strategy"],"falsifier":"Have independent PHITS users attempt the same five workflow classes (input editing, repeated runs, optimization, compilation, post-processing and interpretation) using only the shipped resource packs and rules, with no author coaching, and measure whether the agents complete the tasks without uncaught errors or human rescue.","tokens_in":2955,"feed_emoji":"⚛️","tokens_out":580,"duration_ms":6753,"temperature":0.7,"pith_summary":"Monte Carlo particle transport codes like PHITS are powerful but demand deep expertise in input preparation, execution, error handling, and result analysis. This paper argues that the code itself can close that gap by shipping two complementary AI-ready resource sets: a retrieval-augmented knowledge base built from manuals, lectures, samples, and developer cautions, and a compact agent reference plus PHITS-specific policies and execution rules. Loaded into existing tools such as NotebookLM, Codex, and Claude Code, those resources let general-purpose AI assistants and agents edit inputs, run calculations, inspect errors, analyze results, and even help modify and compile source code across multi-step workflows. Five demonstration tasks spanning input changes, repeated simulations, parameter optimization, compilation, post-processing, and interpretation showed that the agents could complete complex PHITS work when the resources and rules were present. The practical payoff is that particle-transport software can support AI-agent-driven use without building a dedicated code-specific AI application.","feed_headline":"Ship two AI resource packs, and general agents can run PHITS","feed_subtitle":"RAG knowledge base plus agent rules let existing tools edit, execute, and interpret particle-transport workflows","key_machinery":"Two complementary AI-ready resource sets: a bundled knowledge base for RAG-based assistants and a compact agent reference combined with PHITS-specific policies and execution rules that together let existing AI tools drive the full workflow.","core_discovery":"When PHITS is shipped with a RAG knowledge base and a compact agent reference plus PHITS-specific policies and execution rules, general-purpose AI assistants and agents can edit inputs, execute calculations, inspect errors, analyze results, and assist with source modification and compilation across complex multi-step workflows, without a dedicated code-specific AI application.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Ship two AI packs so general agents run full PHITS workflows","RAG base plus agent rules let tools edit execute and interpret PHITS","Bundle knowledge and agent refs for AI-driven particle transport","AI agents handle PHITS inputs runs and analysis with two resource packs","Ship RAG and policies so Codex Claude drive PHITS multi-step tasks"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That success on five author-curated demonstration tasks with NotebookLM, Codex, and Claude Code is representative enough of real PHITS user work to support the broader claim that bundling such resources enables reliable AI-agent-driven particle-transport workflows in practice.","fun_headline_variants_meta":{"raw":{"variants":["Ship two AI packs so general agents run full PHITS workflows","RAG base plus agent rules let tools edit execute and interpret PHITS","Bundle knowledge and agent refs for AI-driven particle transport","AI agents handle PHITS inputs runs and analysis with two resource packs","Ship RAG and policies so Codex Claude drive PHITS multi-step tasks"]},"model":"grok-4.5","effort":"low","cost_usd":0.004792,"raw_usage":{"total_tokens":1326,"prompt_tokens":791,"num_sources_used":0,"completion_tokens":72,"cost_in_usd_ticks":47920000,"prompt_tokens_details":{"text_tokens":791,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":463,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":791,"tokens_out":72,"duration_ms":3943,"temperature":1.0,"reasoning_tokens":463,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T02:01:55.077082+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Have independent PHITS users attempt the same five workflow classes (input editing, repeated runs, optimization, compilation, post-processing and interpretation) using only the shipped resource packs and rules, with no author coaching, and measure whether the agents complete the tasks without uncaught errors or human rescue.","supporting_citations":[],"review_version":1}