{"id":"983ad1bb-cce6-4018-85b1-e8e19791f077","arxiv_id":"2605.26646","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"UnityMAS-O is a general RL optimization framework for LLM-based multi-agent systems that treats complete workflows as the optimization unit and supports role-specific credit assignment via a Ray-based runtime extending verl.","lead":"UnityMAS-O is a framework that lets users optimize entire LLM-based multi-agent workflows with reinforcement learning instead of manual prompts and rules. If it works as described, it could reduce the engineering effort needed to train coordinated agent teams for tasks like search and code generation.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption directly identifies the extensibility claim that underpins the strongest_claim. Because the supplied abstract already states the design intent and reports successful instantiations on diverse tasks, the full manuscript would be needed only to verify implementation details, not to surface a new load-bearing flaw. The current verdict therefore requires no adjustment.","tokens_in":1816,"tokens_out":268,"duration_ms":19796,"concrete_test":"Reproduce one of the three reported workflows (e.g., the reflective code-generation task) using only the four object definitions and the published Ray controller interface; confirm that no modifications to the verl PPO loop or star-topology scheduler were required.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract presents UnityMAS-O as a framework built on four first-class objects plus a Ray star-topology extension to verl, with explicit claims that users can instantiate new workflows, roles, rewards, and mappings without altering core optimization code. The three reported instantiations (retrieval-augmented QA, iterative search, reflective code generation) demonstrate measurable improvement after RL optimization on the stated benchmarks. No internal contradiction, missing derivation, or unsupported scaling claim appears in the provided description that would undermine the reusable-substrate assertion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents UnityMAS-O, a general RL optimization framework for LLM-based multi-agent systems. It models complete workflows as the optimization unit using four first-class objects (logical agent roles, graph trajectories, user-defined rewards, and agent-model mappings) that decouple logical agents from physical model parameters and support flexible credit assignment and parameter sharing. The framework extends verl with a Ray-based star-topology runtime separating a central controller from model-local workers. Three instantiations are described (retrieval-augmented QA, iterative agentic search, reflective code generation) with the claim that multi-agent RL yields improvements over manually specified workflows on Natural Questions, HotpotQA, and held-out code tasks, especially for smaller models and strict code metrics, positioning UnityMAS-O as a reusable substrate for trainable multi-agent RL systems.","tokens_in":1894,"tokens_out":366,"duration_ms":26918,"significance":"If the experimental results hold with proper validation, the work would be significant for providing the first unified RL interface that treats multi-agent workflows as first-class optimization units rather than single-policy trajectories. The abstractions for role-specific rewards, graph trajectories, and configurable model sharing address a clear gap in existing single-agent RL post-training frameworks and could enable broader adoption of multi-agent RL for LLM systems.","major_comments":[{"comment":"Abstract: the central claim that 'multi-agent RL improves manually specified workflows after optimization, with especially large gains for smaller models and strict code all-passed metrics' is asserted without any quantitative results, baselines, error bars, statistical tests, or experimental details. This directly undermines evaluation of whether the three instantiations support the reusable-substrate conclusion.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and positive assessment of the work's potential significance. We address the major comment on the abstract below and will revise the manuscript to strengthen the presentation of results.","responses":[{"response":"We agree that the abstract as written makes a qualitative claim without supporting quantitative details, which limits the ability to evaluate the strength of the evidence. The full paper contains experimental results across the three instantiations (retrieval-augmented QA, iterative agentic search, and reflective code generation) on Natural Questions, HotpotQA, and held-out code tasks, including comparisons to manually specified workflows. In the revised version, we will update the abstract to include specific quantitative highlights (e.g., relative improvements, particularly for smaller models and strict metrics) drawn from the results section, while keeping the abstract concise. We will also ensure the abstract references the experimental setup at a high level.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the central claim that 'multi-agent RL improves manually specified workflows after optimization, with especially large gains for smaller models and strict code all-passed metrics' is asserted without any quantitative results, baselines, error bars, statistical tests, or experimental details. This directly undermines evaluation of whether the three instantiations support the reusable-substrate conclusion."}],"tokens_in":1459,"tokens_out":290,"duration_ms":16668,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"Colleague,\n\nThe main point is that UnityMAS-O extends verl with four first-class objects—logical roles, graph trajectories, user-defined rewards, and agent-model mappings—plus a central Ray controller for workflow execution and distributed PPO updates. This lets the entire multi-agent interaction be the optimization target instead of single policies, and users can add new roles or rewards without touching the training loop.\n\nThe design is straightforward and addresses a clear gap: most current RL post-training tools are single-agent only. The three instantiations on retrieval QA, iterative search, and reflective code generation show how the abstractions map to real tasks, and the claim of larger gains for smaller models and strict code metrics is at least directionally plausible.\n\nThe obvious weakness is that the abstract asserts measurable improvement on Natural Questions, HotpotQA, and held-out code tasks but gives no numbers, baselines, variance, or even basic experimental settings. Without those, it is impossible to tell whether the framework actually delivers usable gains or whether the runtime overhead is manageable. The soundness therefore rests entirely on the full paper’s experiments, which are not visible here.\n\nThis work is aimed at people already running verl or similar libraries who want to move multi-agent LLM pipelines from prompt engineering to end-to-end RL. A practitioner who needs the exact abstractions described would get immediate value from the code structure; a theorist looking for new algorithms would not.\n\nIt deserves peer review. The systems contribution is concrete and the gap it fills is real, even if the current evidence is thin. A referee can check whether the reported improvements hold up and whether the user extensibility claim survives actual use.","headline":"UnityMAS-O is a practical verl extension adding multi-agent abstractions and a Ray runtime so workflows can be treated as RL units, but the abstract supplies zero numbers or baselines to judge the claimed gains.","tokens_in":2423,"tokens_out":418,"would_cite":false,"duration_ms":41419,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"UnityMAS-O turns manually orchestrated LLM multi-agent workflows into trainable multi-agent RL systems by treating the full workflow as the optimization unit.","keywords":["LLM multi-agent systems","multi-agent reinforcement learning","workflow optimization","role-based credit assignment","parameter sharing","distributed RL runtime","PPO for agents"],"falsifier":"A user defines a new multi-agent workflow (for example, a negotiation task) using only the four objects and the provided runtime, then checks whether the optimization loop runs without any changes to the training or buffering code.","tokens_in":2726,"feed_emoji":"🤖","tokens_out":679,"duration_ms":25331,"temperature":0.7,"pith_summary":"Most LLM-based multi-agent systems rely on hand-crafted prompts and rules, with little unified reinforcement learning to improve the agents. UnityMAS-O supplies four first-class objects—logical agent roles, graph trajectories, user-defined rewards, and agent-model mappings—to represent and optimize an entire workflow at once. These objects allow role-specific credit assignment and any combination of parameter sharing or separation across agents. A Ray-based star-topology runtime keeps workflow execution separate from distributed PPO-style model updates, so users can add new agents or rewards without altering the core training loop. Experiments on retrieval-augmented QA, iterative search, and code generation show measurable gains after optimization, especially for smaller models.","feed_headline":"Framework converts LLM multi-agent workflows into RL systems","feed_subtitle":"Four abstractions let users optimize entire workflows with role-level rewards and flexible model sharing, yielding gains on QA and code task","key_machinery":"Four first-class objects (logical agent roles, graph trajectories, user-defined rewards, agent-model mappings) together with a Ray-based star-topology runtime that separates workflow control from rollout and distributed updates.","core_discovery":"By representing any LLM multi-agent workflow through logical agent roles, graph trajectories, user-defined rewards assigned at role/turn/trajectory levels, and flexible agent-model mappings, UnityMAS-O converts the workflow into a single optimization unit that supports full, partial, or no parameter sharing and delivers multi-agent RL training via a central controller and model-local workers.","pith_inferences":["The same abstractions could let teams swap in different agent roles or reward functions on existing workflows with minimal additional engineering.","Because the runtime already records structured trajectories, the framework could support credit-assignment methods beyond PPO without changing the controller.","The decoupling of logical roles from physical models may make it easier to test heterogeneous model sizes within one multi-agent system."],"forward_implications":["Multi-agent RL improves performance over the original manually specified workflows on Natural Questions, HotpotQA, and code tasks.","Smaller models receive especially large gains, including on strict all-passed code metrics.","Rewards and credit can be assigned independently at role, turn, or full-trajectory levels.","The same infrastructure supports full sharing, full separation, or partial sharing of model parameters across logical agents."],"fun_headline_variants":["UnityMAS-O RL-optimizes multi-agent LLM workflows","Workflows as RL units with role-level rewards","Flexible model sharing for multi-agent LLM RL","Role turn trajectory rewards in LLM agent RL"],"cache_read_input_tokens":64,"weakest_assumption_plain":"That the four first-class objects and the Ray-based runtime can be implemented and used by practitioners without rewriting core optimization infrastructure.","fun_headline_variants_meta":{"raw":{"variants":["UnityMAS-O RL-optimizes multi-agent LLM workflows","Workflows as RL units with role-level rewards","Flexible model sharing for multi-agent LLM RL","Role turn trajectory rewards in LLM agent RL"]},"model":"grok-4.3","cost_usd":0.005472,"raw_usage":{"total_tokens":2590,"prompt_tokens":749,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":54715500,"prompt_tokens_details":{"text_tokens":749,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1784,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":749,"tokens_out":57,"duration_ms":8583,"temperature":1.0,"reasoning_tokens":1784,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:14:32.904464+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A user defines a new multi-agent workflow (for example, a negotiation task) using only the four objects and the provided runtime, then checks whether the optimization loop runs without any changes to the training or buffering code.","supporting_citations":[],"review_version":1}