{"id":"28aa9e53-1d0f-4cc1-8aa2-315549b71558","arxiv_id":"2607.27177","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CE-CM infers discrete task-invariant partner capability vectors online via approximate Bayesian simulate-and-compare, and CE-CM-Div improves estimates when humans use diverse suboptimal strategies.","lead":"The paper shows an agent can learn what a new teammate can and cannot do from a few shared tasks, then reuse that model on later tasks without pre-training on partner populations. That matters for household robots and other open-world collaborators who must adapt to unknown humans rather than fixed game partners.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Offline human evaluation does not establish that better capability estimates improve closed-loop teaming.","rationale":"The reader correctly flags the Section 3.3 alignment assumption and known capability-to-transition structure as external-validity limits, and already notes the offline-only human evidence. The more load-bearing gap for the packaged strongest claim is that CE-CM-Div’s human result is purely an estimation metric (Hamming / posterior size) while the paper’s own Overcooked simulation shows estimation gains need not translate into coordination gains when multiple strategies are feasible. That makes the abstract/conclusion leap from “better capability estimates on human data” to “essential for robust human–AI teaming” unsupported by a closed-loop human test. Verdict remains CONDITIONAL (same bucket as the reader) but the condition should explicitly include a closed-loop human coordination result, not only the alignment/structure assumptions. I partially agree with the reader: their weakest assumption is valid and limits validity, yet it is not the single softest point under the strongest claim as stated; the offline-to-teaming gap is.","tokens_in":25967,"tokens_out":593,"duration_ms":12696,"concrete_test":"Re-run the 15-participant Overcooked protocol with online CE-CM-Div: after each of the first k tasks, update β and ˆc*, replan the next task with M_ˆc*, and log correction rate, unproductive-action ratio, and delivery success versus a no-update optimistic baseline (and versus CE-CM). If Hamming improves as in Fig. 13 but correction/success metrics do not improve significantly over baseline, the teaming half of the strongest claim does not hold for humans.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim packages three results: (1) CE-CM recovers task-invariant capabilities via ABC simulate-and-compare, (2) those estimates reduce infeasible assignments in simulation, and (3) CE-CM-Div substantially improves capability estimates on human Overcooked data. Claim (3) is only offline Hamming-distance recovery on 225 recorded trajectories (Section 5.3); the paper never closes the loop by using CE-CM-Div estimates to replan with the same humans and measure corrections, plan overlap, unproductive actions, or task success. Section 5.2 already shows that better capability estimates need not improve coordination under strategy underspecification (corrections stay high for Types 1–2 despite Hamming drop and fewer unproductive actions). The human study therefore supports only the intermediate claim that diversity-aware ABC matches human trajectories better, not the abstract’s framing that this yields more robust human–AI teaming. The reader’s alignment/known-structure assumption is real but secondary: even under perfect alignment and known gates, the human evidence stops at estimation accuracy.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper reframes multi-task ad-hoc teamwork as joint planning with decentralised execution under hidden, task-invariant partner capabilities in a contextual MMDP. It introduces CE-CM, an ABC-style approximate Bayesian method that samples capability vectors, simulates joint trajectories under a planner, and updates an independent-Bernoulli belief from accepted samples; the MAP estimate then induces a capability-conditioned model for planning the next task. CE-CM-Div extends the likelihood comparison to a set of diverse planner rollouts to handle suboptimal or multi-modal partner behaviour. Simulated results in TidyUP and Overcooked show rapid recovery of ground-truth capability vectors, fewer infeasible partner assignments, and adaptation when capabilities change. An offline study of 225 Overcooked trajectories from 15 humans shows CE-CM-Div substantially lowers Hamming error relative to single-trajectory CE-CM.","tokens_in":26209,"tokens_out":1513,"duration_ms":38402,"significance":"If the results hold, the work offers a clear, interpretable alternative to population-trained AHT policies: a reusable partner model that transfers across tasks without pre-training on partner populations. The capability-as-transition-constraint formulation, the online ABC loop, and the diversity extension are concrete methodological contributions, and the paper ships public code plus a human trajectory dataset. The TidyUP results and the feasibility/safety gains in Overcooked are convincing evidence that capability estimates can improve action allocation when behaviour is largely feasibility-driven. The human study usefully documents that planner-optimal trajectories poorly match people, motivating diversity-aware inference. These strengths make the paper a solid contribution to human–AI teaming and multi-task AHT, provided the claims about closed-loop human teaming are aligned with the evidence.","major_comments":[{"comment":"Abstract and §5.3/§7 frame CE-CM-Div as essential for robust human–AI teaming, but the human evaluation is purely offline Hamming-distance recovery on recorded trajectories (225 episodes, fixed assistant policy). There is no closed-loop experiment in which CE-CM-Div estimates are used to replan with the same participants and measure corrections, plan overlap, unproductive actions, or task success. §5.2 already shows that better capability estimates need not reduce corrections under strategy underspecification (Types 1–2). The human result therefore supports only improved estimation under behavioural diversity, not the stronger teaming claim. Either add a closed-loop human evaluation or revise the abstract/conclusion to match the intermediate claim actually tested.","section":"Abstract; §5.3; §7"},{"comment":"§3.3 assumes observed trajectories are feasible under the partner’s true CMMDP because Ag2 can correct Ag1 and enforce the aligned joint plan. This alignment channel is load-bearing for the ABC likelihood: without it, ego misallocation or simultaneous capability errors would corrupt τ_obs and the simulate-and-compare update would not identify c. The paper notes the assumption but does not stress-test it (e.g., noisy/partial corrections, delayed corrections, or no correction channel). A sensitivity experiment or a clearly scoped limitation stating that identification holds only under aligned execution is needed for the central inference claim.","section":"§3.3"},{"comment":"The method assumes the capability-to-transition gating structure is known a priori (which bits enable which transitions; §3.2 and the f pruning function in App. B.2). Combined with a hand-specified discrete capability vocabulary, this weakens the “task-agnostic / no pre-coordination” framing relative to methods that learn latent partner structure from data. The paper should state explicitly what must be known in advance versus what is inferred online, and discuss how misspecified gates would affect posterior recovery.","section":"§3.2; Appendix B.2"},{"comment":"Baselines are limited to optimistic (and, in TidyUP, pessimistic) non-adaptive models. There is no comparison to type-based AHT, latent-conditioned policies, behaviour cloning of the partner, or preference-learning approaches discussed in §2. Without at least one adaptive partner-modelling baseline on the same multi-task protocol, it is hard to judge whether capability vectors are competitive as a representation rather than merely better than assuming full capability. Adding one such baseline in simulation would substantially strengthen H1–H2.","section":"§5.1.1; §5.2.1"}],"minor_comments":[{"comment":"Eq. (1) minimises a sum of trajectory losses over G_seen, but the implemented estimator is an ABC posterior mean/threshold on independent Bernoullis (§4.1). Briefly state that the MAP of the ABC approximation is treated as a surrogate for (1).","section":"§3.3; §4.1"},{"comment":"Independent Bernoulli factorisation of the posterior (§4.1) ignores correlations among capabilities (e.g., room-linked pick/place bits in TidyUP). A short note on when this approximation fails would help.","section":"§4.1"},{"comment":"Figure 1 and Algorithm 1 are clear; however, the acceptance threshold ε, prior P(c_i=1)=0.8, ψ=0.5, and δ=0.35 are scattered across Appendix B. A small hyperparameter table in the main text (or early appendix pointer) would aid reproducibility.","section":"Appendix B"},{"comment":"In Overcooked, cosine similarity with ε=0.03 on flattened states is quite tight; a brief justification or sensitivity check (analogous to App. C.3 for δ) would be useful.","section":"Appendix B.2"},{"comment":"Typos/consistency: “CApability Modelling” in the CAMO paragraph; duplicate “Transitions are deterministic…” block in App. A.2; “place_study / pick_study / move_study” descriptions say “kitchen” in Table A.1.","section":"§1; Appendix A"},{"comment":"Related work on ZSC/AHT is solid; a short pointer to recent assistance games / theory-of-mind planning work beyond the cited goal-recognition papers would round out §2.2.","section":"§2"}],"recommendation":"major_revision","confidential_remarks":"The technical core (multi-task capability ABC + diversity extension) is publishable and the simulation evidence is honest about the coordination gap in Overcooked. The main risk is overclaim in the abstract relative to an offline-only human study. If the authors narrow the human claim and clarify the alignment/known-gate assumptions, this is appropriate for a solid methods venue; I would not reject on novelty grounds. Fit is good for an AI/HRI journal that values empirical multi-agent collaboration."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core here is reframing multi-task ad-hoc teamwork as joint planning under hidden, task-invariant partner capabilities, then doing online ABC-style simulate-and-compare (CE-CM) with no partner-population pre-training. That is a real step past the usual single-task robust-policy AHT line. CE-CM-Div’s set-based likelihood is a modest but practical fix when humans do not match one planner trajectory.\n\nWhat they do well is match claims to evidence and stay honest about the negative. In TidyUP, Hamming distance drops fast, corrections fall, plan IoU rises versus optimistic and pessimistic baselines. In Overcooked they recover approximate capability vectors and cut unproductive assignments, then openly report that corrections and coordination do not fully improve under strategy underspecification—capabilities constrain feasibility, they do not pick the convention. The 225-trajectory human set (15 people) shows CE-CM-Div accepts far more samples and cuts Hamming error versus single-trajectory CE-CM. Code and hyperparameter appendices are above average for this area. Self-cite to CAMO is the inference template; the multi-task discrete setting and decentralized planning loop are new enough.\n\nSoft spots in proportion: the load-bearing alignment assumption (partner corrects the ego agent so observed trajectories stay feasible under true c) and the known capability-to-transition gates limit external validity—those are real, not fatal for a methods paper. Free parameters (ε, ψ, N, prior, δ, k, planner knobs) are normal for ABC/MCTS work but need sensitivity care. The stress-test lands: the human study is offline estimation only. They never close the loop by replanning with the same people and measuring corrections, success, or unproductive actions. Section 5.2 already showed better Hamming need not mean better coordination, so the abstract’s “robust human–AI teaming” framing overreaches the offline Hamming result. Still an intermediate result worth having.\n\nThis is for people in AHT, HRI partner modeling, and capability/preference layered models. Math is standard ABC + CMMDP, citations are fair, data support the stated claims. I would send it to referees; engage if you care about transferable partner models rather than another population-trained ego policy.","headline":"Solid multi-task AHT methods paper: online capability inference without population training works in sim, diversity helps offline human matching, but closed-loop teaming gains are not shown.","tokens_in":26884,"tokens_out":571,"would_cite":true,"duration_ms":18572,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Hidden partner abilities can be inferred from a few joint tasks and reused to plan safer teamwork across new tasks.","keywords":["ad-hoc teamwork","human–agent interaction","behaviour modelling","capability estimation","contextual multi-agent MDP","approximate Bayesian computation","diverse planning"],"falsifier":"If, with the same known capability-to-transition structure and correction channel disabled or noisy, CE-CM and CE-CM-Div fail to drive Hamming distance toward zero over a few tasks on held-out partners, or CE-CM-Div no longer beats single-trajectory CE-CM on human Overcooked trajectories, the central claim fails.","tokens_in":26811,"feed_emoji":"🤝","tokens_out":927,"duration_ms":21412,"temperature":0.7,"pith_summary":"Most ad-hoc teamwork methods train a single ego policy for one fixed task and treat the partner as a black box. This paper instead treats repeated collaboration as joint planning under hidden, task-invariant partner capabilities—what the partner can and cannot do. CE-CM samples candidate capability vectors, simulates the joint trajectories each would produce, and keeps those that match what was observed, building an approximate posterior without population pre-training. The resulting estimate induces a capability-conditioned multi-agent model so the agent can plan the next task while executing only its own actions. When humans choose among many valid strategies, CE-CM-Div scores each hypothesis against diverse planner rollouts rather than one optimal path. In a household planning domain and in Overcooked, the method recovers capabilities, cuts infeasible assignments, and adapts when abilities change; on 225 human trajectories it needs diversity-aware matching to get reliable estimates.","feed_headline":"Few joint tasks reveal what a partner can and cannot do","feed_subtitle":"Capability vectors transfer across tasks and cut infeasible assignments; humans need diversity-aware matching","key_machinery":"CE-CM (Capability Estimation via Contextual Models): Approximate Bayesian Computation over discrete capability vectors—sample candidates, joint-plan under each, accept those whose simulated trajectories fall within a distance threshold of the observation, aggregate into a Bernoulli belief, and plan on the induced capability-conditioned CMMDP; CE-CM-Div replaces the single rollout with a diverse trajectory set per hypothesis.","core_discovery":"Partner capabilities—latent binary constraints on which joint transitions are feasible—are a reusable, task-agnostic representation that an ego agent can recover online from a handful of observed joint trajectories via approximate Bayesian simulate-and-compare, then use to induce contextual multi-agent MDPs for decentralised joint planning; when behaviour is diverse or suboptimal, matching against sets of diverse rollouts (CE-CM-Div) is necessary for accurate inference, especially with humans.","pith_inferences":["The same simulate-and-compare loop could be tried for continuous or graded ability parameters if the discrete bit vector is too coarse for physical skills.","Without a known capability-to-transition map, the method would need to jointly discover structure and values—closer to structure learning than pure ABC filtering.","Scaling beyond two agents will force factorised or role-based planning, or the joint planner cost will dominate any inference gain.","Correction-based alignment during data collection may understate how hard inference is when both agents act without a referee."],"forward_implications":["Agents can accumulate an explicit, interpretable model of one partner across chores or layouts without retraining a population policy.","Feasibility-aware planning reduces assignments of actions the partner cannot execute, improving safety even when full behavioural prediction remains ambiguous.","When multiple strategies are equally valid, diversity-aware likelihoods are required for capability inference from human data.","Capability estimates can be updated online if the partner’s abilities change mid-collaboration.","Capability models alone do not resolve preference or convention ambiguity; joint capability-plus-preference models become the natural next representation."],"fun_headline_variants":["Few joint tasks expose a partner's hidden capability limits","Online Bayesian sampling recovers reusable partner capability vectors","Capability estimates from sparse trajectories cut infeasible joint actions","Diverse rollouts sharpen capability inference for suboptimal human partners","Task-agnostic capability vectors enable decentralised multi-task teamwork"],"cache_read_input_tokens":23936,"weakest_assumption_plain":"Observed joint trajectories are assumed to already be feasible under the partner’s true abilities because the partner can correct the ego agent and enforce the aligned plan, and the mapping from capability bits to which transitions they gate is known in advance.","fun_headline_variants_meta":{"raw":{"variants":["Few joint tasks expose a partner's hidden capability limits","Online Bayesian sampling recovers reusable partner capability vectors","Capability estimates from sparse trajectories cut infeasible joint actions","Diverse rollouts sharpen capability inference for suboptimal human partners","Task-agnostic capability vectors enable decentralised multi-task teamwork"]},"model":"grok-4.5","effort":"low","cost_usd":0.004022,"raw_usage":{"total_tokens":1282,"prompt_tokens":863,"num_sources_used":0,"completion_tokens":80,"cost_in_usd_ticks":40224000,"prompt_tokens_details":{"text_tokens":863,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":339,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":863,"tokens_out":80,"duration_ms":5558,"temperature":1.0,"reasoning_tokens":339,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T10:53:53.220956+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If, with the same known capability-to-transition structure and correction channel disabled or noisy, CE-CM and CE-CM-Div fail to drive Hamming distance toward zero over a few tasks on held-out partners, or CE-CM-Div no longer beats single-trajectory CE-CM on human Overcooked trajectories, the central claim fails.","supporting_citations":[],"review_version":2}