{"id":"b84a63c6-1500-47d1-806d-ac615d87f745","arxiv_id":"2608.13173","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SkillShapley assigns each instruction step of an LLM agent skill a Shapley value, and its BAES estimator recovers that ranking with fewer skill-variant evaluations than standard Shapley approximations.","lead":"This paper treats each instruction step in an LLM agent skill as a player in a cooperative game and uses Shapley values, approximated by a new sampling method named BAES, to score each step's contribution to task success. The approach ranks steps with fewer skill-variant evaluations than standard Shapley estimators on three SkillsBench skills, giving skill authors a signal for which steps to prune or reinforce.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical target is defined by just three benchmark instances per coalition (Appendix A: v in {0,1/3,2/3,1}); exact Shapley, BAES error, and removal validation all depend on this coarse reference, with no stability analysis. Rankings could be an artifact of the particular three-instance subset.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing point, and I agree with it. The paper makes two linked claims: (1) Shapley values are behaviorally meaningful for skill-step attribution, and (2) BAES approximates them efficiently. Both are tested only against exact Shapley computed from v(S) on three fixed benchmark instances. The efficiency claim is relative to a reference that may itself be unrepresentative; the meaningfulness claim's removal validation uses the same v(S) to both rank steps and measure the consequence of removal, so it is not fully independent evidence. Comparing Shapley against Individual and LOO on the same v is still a reasonable relative check, so this is not a fatal flaw by itself, but without an independent instance sample the general conclusion is not established. The prose is honest about limitations, and the mathematics of the Shapley formulation and the BAES procedure are standard and clearly presented. The proposed test is feasible because the exact coalition lattices are only 512-2048 configurations, and it would directly settle whether the three-instance reference changes the ranking, the removal curve, and the BAES error comparison. Therefore the conditional verdict should remain unchanged rather than moving to accept or reject.","tokens_in":10620,"tokens_out":7113,"duration_ms":70696,"concrete_test":"For each of the three SkillsBench task-skill pairs, evaluate all coalitions in the exact lattice (or a stratified sample of at least 200 coalitions) on a new independent set of 30 benchmark instances from the same task distribution; compute v'(S), re-derive exact Shapley, and rerun BAES at the same unique-configuration budgets. Check: (1) Spearman rank correlation between the 3-instance and 30-instance Shapley step rankings; (2) whether the top-3 removal drop in Figure 2 remains above the baselines under v'; (3) whether BAES still has lower error than MC/truncation baselines when error is measured against the 30-instance exact Shapley. If rankings shift materially or the removal advantage disappears, the three-instance reference is load-bearing and the claims need re-qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5 and Appendix A evaluate every coalition on exactly three benchmark instances, so v(S) (Eq. 3) takes values in {0, 1/3, 2/3, 1} and g = 1/3. The 'exact Shapley' used as ground truth in RQ1 and RQ2 is exact only for this three-instance aggregate; it is a low-resolution, potentially unrepresentative estimate of the Shapley value under the benchmark distribution. This matters for both central claims. First, the behavioral validation in Figure 2 removes top-ranked steps and measures average v over remaining coalitions using the same v(S) that generated the Shapley ranking, so the removal curve is not an independent confirmation; it reuses the target being validated. Second, BAES's 'lower approximation error' is measured against this same three-instance reference; if the reference is noisy or biased, BAES can appear to approximate a particular instance realization rather than the true skill-step value. The paper acknowledges in Section 4.2 that rewards are 'discrete and noisy under limited repeated evaluation,' and the Conclusion lists limitations, but no bootstrap, instance-subsample sensitivity, or independent-instance check is provided. Without such a check, the differentiated Shapley rankings and the BAES efficiency claim are not established beyond the three chosen instances.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formulates step-level attribution in LLM agent skills as a Shapley-value computation over skill steps, treating each instruction block as a player and benchmark success rate as the utility function. It proposes SkillShapley, which uses a two-phase budgeted estimator called BAES: a warmup stage that builds broad cache coverage over coalition sizes, followed by an adaptive stage that acquires coalitions maximizing reusable one-flip marginal edges. The empirical evaluation uses three SkillsBench skills (offer-letter-generator, manufacturing-fjsp-optimization, dialogue parser) with exact Shapley enumeration as the reference, reports top-removal validation curves, and compares BAES against Monte Carlo, Quasi-Monte Carlo, paired Monte Carlo, and size-k-truncated Shapley estimators under matched unique-configuration budgets. The central claims are that Shapley values are behaviorally meaningful for skill steps and that BAES approximates them more efficiently than existing estimators.","tokens_in":10916,"tokens_out":4187,"duration_ms":36135,"significance":"If the empirical claims are robust, the framework addresses a genuine gap: existing skill optimization methods treat skills as whole units, whereas SkillShapley provides per-step valuation with a practical budgeted estimator. The matched-budget comparison is a sound experimental design, and the cache-aware edge-reuse mechanism is a plausible way to reduce the number of expensive agent executions. The paper also gives credit to the problem by explicitly isolating the approximation question from the cost model. However, the evidence base is narrow: only three skills, each with a utility measured on three benchmark instances, and no stability analysis. The significance is therefore conditional on whether the coarse utility function is representative of the benchmark distribution.","major_comments":[{"comment":"The utility v(S) is computed on exactly three benchmark instances, so v(S) takes values in {0, 1/3, 2/3, 1} and g = 1/3. This coarse, instance-specific realization is the 'exact Shapley' reference used in Section 5.2 and Figure 3. The paper provides no bootstrap, instance-subsample, or otherwise randomized stability analysis, despite acknowledging in Section 4.2 that rewards are 'discrete and noisy under limited repeated evaluation.' If the specific three-instance subset is not representative, the Shapley rankings, the removal curves, and the BAES approximation-error comparisons are all conditioned on that particular subset. This is load-bearing for both research questions, and the manuscript should add a stability analysis, for example by subsampling or bootstrapping over instances for at least one skill and reporting the sensitivity of the Shapley ranking and of BAES's relative error.","section":"Appendix A; Eq. (3)"},{"comment":"The removal validation in Figure 2 is not an independent confirmation of the Shapley rankings. The same v(S) from Eq. (3) is used both to compute the exact Shapley ranking and to measure the post-removal success rate (the mean of v over remaining coalitions). Under this protocol, a ranking that is overfit to the three-instance realization will automatically show a sharper drop than a more generalizable ranking. The validity claim would be substantially strengthened if the removal curves were evaluated on held-out benchmark instances that were not used to estimate the Shapley values, or at least on a bootstrap of the instances.","section":"Figure 2"},{"comment":"The stated research question RQ2 asks whether BAES can approximate exact Shapley rankings under a small budget, but the experiment in Figure 3 and the metrics in Appendix A report only player-wise MAE against the exact Shapley vector. No rank-oriented metric such as Kendall's tau, Spearman correlation, or top-k step overlap is reported. Since the paper's practical motivation is identifying high- and low-value steps rather than reporting numerically accurate Shapley scalars, the approximation claim would be better supported by reporting ranking-quality metrics alongside value error, particularly under the small budgets that are the intended use case.","section":"Section 5.3; RQ2"}],"minor_comments":[{"comment":"There is a typo in the abstract: 'i.e.,there' should be 'i.e., there'.","section":"Abstract"},{"comment":"The caption states the y-axis is 'error relative to full Shapley' but does not specify the metric; Appendix A says player-wise MAE is used. Please state that explicitly in the caption for self-containedness.","section":"Figure 3 caption"},{"comment":"The paper states that all model calls use temperature T = 0 but does not mention whether decoding is fully deterministic; if it is, the only source of stochasticity in v(S) is benchmark-instance selection, which reinforces the need for an instance-stability analysis in the major comments.","section":"Section 5.1"},{"comment":"The paper does not state data or code availability; adding a reproducibility statement would improve the practical value of the method for practitioners who might want to apply SkillShapley to their own skills.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript depends heavily on very recent or non-archival references for its empirical target (SkillsBench [11]) and for several related methods; the editor may wish to verify the verifiability of these benchmarks. The main correctness risk is the use of a three-instance utility as the 'exact' reference, which I have flagged in the major comments; it is fixable within the manuscript's scope by adding stability analysis and independent removal validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper is the first to my knowledge to treat individual steps in an LLM agent skill as Shapley players over retained-subset coalitions, and BAES, its cache-aware active sampling, is a concrete and sensible algorithmic proposal. Second, the empirical support is thin: every coalition payoff is the average over exactly three benchmark instances, so v(S) is always 0, 1/3, 2/3 or 1. There is no bootstrap, no instance-subsample sensitivity, and no indication of how stable the rankings are. The stress-test note is right: the exact-Shapley 'ground truth' is exact only for that three-instance aggregate, so both the approximation error and the removal curves inherit its noise.\n\nWhat the paper does well: the problem formulation is clean; the size-stratified Shapley decomposition is standard but appropriate; and the matched-budget comparison against exact Shapley is the correct experimental design for isolating sampling policy from raw API cost. The paper also deserves credit for being candid: it explicitly says BAES is a biased approximation for ranking recovery and lists workflow coupling as a limitation. The case studies give concrete examples of high- and low-value steps that look plausible.\n\nThe soft spots, in order. The three-instance utility is the load-bearing one. The authors acknowledge in Section 4.2 that rewards are 'discrete and noisy under limited repeated evaluation,' but they never check whether the Shapley rankings, the BAES error curves, or the removal validation change when you regrade on different or additional instances. Without that, the differentiated step values could be an artifact of which three inputs happened to be selected. The removal validation in Figure 2 reuses the same v(S) that generated the Shapley ranking, so it is not an independent behavioral confirmation; it mostly shows Shapley is consistent with its own utility function. That is not fatal, but it should be labeled as a consistency check, not validation. No code or data released, and several hyperparameters are hand-set with no sensitivity analysis.\n\nThe audience is researchers and practitioners working on LLM agent skills; the method is useful as a diagnostic if you have enough evaluation budget. If I were editing, I would send this to review. The core idea is new, the method is specified well enough to reimplement, and the limitations are honestly stated. But I would require a stability analysis and at least some additional evaluation instances before publication. The paper deserves a serious referee's time, not a desk reject.","headline":"Solid Shapley-for-agent-skills idea with a cache-aware sampler, but the three-instance evaluation grid makes all empirical claims fragile.","tokens_in":11461,"tokens_out":4136,"would_cite":false,"duration_ms":32262,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that Shapley values are behaviorally meaningful for agent skill steps and that BAES approximates them efficiently from a small budget of skill evaluations.","keywords":["Shapley value","agent skills","LLM agents","step attribution","active sampling","coalitional game theory","budgeted approximation","SkillsBench"],"falsifier":"Pick one of the three skills, draw several different three-instance benchmark subsets (or enlarge to ten instances) and recompute exact Shapley rankings; if top-step identity or the relative error of BAES changes substantially, the reported attribution signal is an artifact of the instance triple.","tokens_in":10417,"feed_emoji":"🎯","tokens_out":6958,"duration_ms":58269,"temperature":0.7,"pith_summary":"SkillShapley claims that the steps of an LLM agent skill can be valued as players in a cooperative game: each step's worth is its Shapley value, computed from benchmark success rates of skill variants that keep different subsets of steps. The paper argues these value rankings are behaviorally meaningful because deleting Shapley-top-ranked steps degrades task success faster than deleting steps ranked by isolated or leave-one-out scores, which produce many ties. Because exact Shapley requires exponential coalition evaluations, the paper introduces BAES, a budgeted active sampler that reuses cached one-flip comparisons to approximate the same rankings from far fewer unique skill variants. If correct, skill authors can get actionable step values from a few dozen runs instead of enumerating all subsets.","feed_headline":"Budgeted Shapley sampler ranks LLM skill steps from few variants","feed_subtitle":"Shapley rankings beat isolated and leave-one-out scores, and a cached sampler recovers them from few runs.","key_machinery":"BAES is the key mechanism: a cache-aware, budgeted active sampler for the size-stratified Shapley estimator $\\hat\\phi_i = \\frac{1}{n}\\sum_{k=0}^{n-1}\\hat\\mu_{i,k}$, where $\\hat\\mu_{i,k}$ is the empirical mean of observed one-flip marginal effects $\\Delta_i(S)=v(S\\cup\\{i\\})-v(S)$ in stratum $(i,k)$. It scores strata by $a_{i,k}=\\sqrt{\\hat\\sigma^2_{i,k}+\\epsilon}/\\sqrt{m_{i,k}+1}$ so uncertain and under-sampled strata get priority, and acquires the unevaluated coalition maximizing $A(C)=\\sum_{i:C\\triangle\\{i\\}\\in D} a_{i,k}\\cdot b(v(C\\triangle\\{i\\}))$ over cached one-flip neighbors. The warmup stage builds broad stratum coverage with anchors ($\\emptyset$, all players, singletons, $(n-1)$-subsets) and a greedy cache-aware rule; the adaptive stage spends the remaining budget where cached edges reduce uncertainty most. The design exploits the paper's empirical observations that configuration evaluation dominates cost, rewards are discrete and high-variance, and rewards flatten in large coalition regions.","core_discovery":"The central claim is that step-level Shapley values are a behaviorally meaningful attribution for agent skill steps and that the Boundary-Adaptive Edge Shapley (BAES) procedure approximates them efficiently. Treating each instruction block as a player and each retained-subset skill variant as a coalition, the paper defines the value of a step as its Shapley value with respect to the empirical success rate $v(S)$ on a fixed benchmark subset. Exact Shapley rankings, checked on three low-step-count SkillsBench tasks, produce differentiated step values and yield the fastest performance drop when top-ranked steps are removed, unlike Individual and Leave-One-Out scores that tie. BAES first warms up coverage across coalition sizes and then adaptively evaluates new coalitions chosen to form many reusable one-flip marginal edges with the cache, reporting that it reaches lower approximation error under smaller unique-configuration budgets than Monte Carlo, Quasi-Monte Carlo, paired Monte Carlo, and size-truncated Shapley baselines. The paper reads the resulting patterns as guidance for skill pruning, revision, and creation: high-value steps are procedural bridges connecting conditions to executable decisions, while low-value steps are locally correct but action-incomplete.","pith_inferences":["If the three-instance evaluation is noisy, the stability of rankings across benchmark subsamples should be checked; a bootstrap over instances would tell whether BAES's advantage over Monte Carlo is robust to the choice of evaluation subset.","The boundary-adaptive acquisition rule may transfer to other valuation problems where evaluating a candidate is expensive and rewards are discrete, such as data subset valuation for few-shot prompts or test-suite minimization, though the paper does not claim this.","Because BAES is a biased approximation that optimizes ranking recovery, downstream users should re-run the removal curve after an edit to validate the edit rather than trusting a single value vector.","The paper's limitation about strongly coupled assembly-line workflows implies that a practical deployment should first verify that the coalition space is meaningfully executable before interpreting value differences."],"forward_implications":["With a budget of about $3n^2$ unique skill variants, a practitioner can obtain step value rankings that track the exact Shapley ordering closely enough to guide pruning, inspection, and revision.","Deleting top-ranked Shapley steps causes a sharper drop in benchmark success than deleting steps by Individual, Leave-One-Out, or Random Removal, so the rankings carry behavioral meaning.","The same coalition cache can be reused for multiple attribution questions: one evaluation of a variant contributes one-flip edges to several strata, lowering the per-stratum cost.","Attribution patterns across tasks suggest skills should be written as compact, decision-complete guidance units, with procedural bridges made explicit and action-incomplete helper steps removed or templated.","Token cost of a skill variant is not proportional to coalition size; step-level editing is best seen as removing low-value content rather than guaranteed token savings."],"supporting_citations":[{"why":"Defines the Shapley value that Eq. (5) uses as the attribution target, providing the cooperative-game foundation.","marker":"[8]"},{"why":"Supplies the SkillsBench skills, benchmark instances, and deterministic verifiers used for every coalition evaluation.","marker":"[11]"},{"why":"Provides the Monte Carlo Shapley sampling baseline that BAES is compared against under matched budgets.","marker":"[21]"},{"why":"Supplies the Quasi-Monte Carlo and paired Monte Carlo Shapley estimators used as approximation baselines.","marker":"[25]"},{"why":"Motivates Shapley value-based attribution for model explanations, the interpretive frame applied to skill steps.","marker":"[9]"},{"why":"Establishes Shapley values as a valuation method for data, supporting the transfer to valuing skill steps.","marker":"[10]"},{"why":"Provides the LeastCore contribution method used as an additional baseline in the removal validation.","marker":"[26]"}],"fun_headline_variants":["Adaptive Shapley sampler ranks LLM agent skill steps","SkillShapley: Boundary-adaptive step attribution for agent skills","Efficient Shapley values for skill-step quality in LLM agents","Budgeted Shapley approximation identifies critical skill steps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every reported value comes from evaluating coalitions on just three benchmark instances, so the exact-Shapley reference and BAES's error and removal curves all inherit whatever noise those three instances introduce.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive Shapley sampler ranks LLM agent skill steps","SkillShapley: Boundary-adaptive step attribution for agent skills","Efficient Shapley values for skill-step quality in LLM agents","Budgeted Shapley approximation identifies critical skill steps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1428,"prompt_tokens":985,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":601,"completion_tokens_details":{"reasoning_tokens":371}},"tokens_in":601,"tokens_out":443,"duration_ms":3786,"temperature":1.0,"reasoning_tokens":371,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:06:12.462754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pick one of the three skills, draw several different three-instance benchmark subsets (or enlarge to ten instances) and recompute exact Shapley rankings; if top-step identity or the relative error of BAES changes substantially, the reported attribution signal is an artifact of the instance triple.","supporting_citations":[{"cited_title":"SkillsBench: Benchmarking how well agent skills work across diverse tasks, 2026","cited_arxiv_id":null,"evidence_quote":"Supplies the SkillsBench skills, benchmark instances, and deterministic verifiers used for every coalition evaluation."},{"cited_title":"Covert, Scott M","cited_arxiv_id":null,"evidence_quote":"Supplies the Quasi-Monte Carlo and paired Monte Carlo Shapley estimators used as approximation baselines."},{"cited_title":"Data Shapley: Equitable valuation of data for machine learning","cited_arxiv_id":null,"evidence_quote":"Establishes Shapley values as a valuation method for data, supporting the transfer to valuing skill steps."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the LeastCore contribution method used as an additional baseline in the removal validation."}],"review_version":1}