{"id":"86781dc9-368d-428b-89bd-0700bb1bf41f","arxiv_id":"2608.10042","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"UserToolBench evaluates LLMs on personalized tool-use decisions with hidden user profiles and incomplete requests, and finds the best model reaches only 49.36% exact trajectory accuracy.","lead":"UserToolBench, a new benchmark, tests whether AI assistants can make the right tool calls for a specific user when the user's profile is hidden and requests are incomplete. It shows even the best current models get less than half of these personalized tasks exactly right, suggesting personalization evaluation should focus on decisions, not just style.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference-trajectory validity is load-bearing: low exact accuracy measures agreement with an undisclosed LLM planner until a profile-visible oracle or independent reference validation rules out generator-specific planning behavior.","rationale":"UserToolBench is a well-motivated benchmark, and the relaxed-vs-exact gap is genuinely suggestive that generic task completion and profile-conditioned decision alignment are different capabilities. The paper is transparent about its single-reference limitation and provides auxiliary diagnostics, which are appropriate. However, the headline empirical claim is only interpretable as evidence about personalized delegation if y* encodes genuine user preferences rather than one LLM's planning conventions. The undisclosed generator model and the absence of any external validation create a real risk that the evaluation measures agreement with the reference generator's idiosyncratic behavior. A profile-visible oracle baseline directly tests whether p is the operative variable: if explicit profile access does not substantially improve exact accuracy, then the low hidden-profile scores are not attributable to the difficulty of recovering user preferences from history. The profile-visible experiment is cheap, uses the existing benchmark, and would settle whether the central claim holds. This is exactly the weakest assumption identified by the reader, and my analysis agrees with that identification. Because the concern is concrete, addressable, and does not invalidate the benchmark's core design, the appropriate verdict remains conditional on this additional evidence rather than moving to acceptance or rejection.","tokens_in":19181,"tokens_out":4025,"duration_ms":44020,"concrete_test":"Run a profile-visible oracle baseline: take a random sample of 100–200 instances and give each evaluated model the full persona profile (explicit p) together with the history, request, and tool schemas, scoring with the same exact and relaxed metrics. If average exact accuracy stays near 50% rather than rising toward 80–100%, the reference trajectories are not determined by the profile plus context, so the hidden-profile gap cannot be interpreted as personalization-inference failure. Independently, regenerate references for the same sample with a different LLM family (and, where feasible, human planners) and measure pairwise exact-trajectory agreement; low agreement below roughly 60% would confirm reference-generator arbitrariness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that hidden-profile exact accuracy of 49.36% demonstrates failure of personalized delegation—rests on Section 3.2's premise that the reference trajectories y* = f_ref(p, h, q, T) are valid ground truth for \"correct\" personalized decisions. This premise is not independently secured. The reference generator is described only as \"the planner component of the assistant pipeline\" (Appendix E), with no generator identity disclosed; the human verification interface (Appendix E.1) checks internal consistency, tool validity, and persona consistency, not agreement with real user preferences or with independently elicited expert decisions. The personas themselves are LLM-assisted abstractions from \"privacy-sanitized real interaction traces,\" and no external validation ties them to the originating users or to real-world preference distributions. Consequently, exact matching may reward reproducing the reference generator's planning conventions rather than recovering the user's preferences. The paper's own single-reference caveat (Section 3.4) concedes that exact matching can penalize equally valid preference-compatible trajectories; the severity diagnostics in Appendix D.2 mitigate but do not eliminate this risk because every comparison is against the same reference trajectory. If the reference generator is idiosyncratic in tool ordering, argument grounding, or clarification thresholds, both the absolute 49.36% figure and the exact-vs-relaxed gap would overstate personalization failure. This is the weakest link because it is the bridge from an internal consistency score to a claim about user-aligned decision making.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces UserToolBench, a benchmark for evaluating personalized decision-making in tool-use LLMs under a profile-hidden protocol. Reference trajectories are constructed with access to persistent user profiles, a persona-conditioned user simulator, and an unspecified planner; evaluated models receive only interaction history, the current request, and tool schemas. The benchmark includes 10 profiles, 36 toolsets, 1,065 turns, 170 unique tools, and 799 task instances. Experiments with nine models report best average exact accuracy of 49.36% and a large gap between relaxed task completion and exact trajectory matching, which the authors interpret as evidence that executable task completion does not imply personalized decision alignment. The paper also includes trajectory-diversity diagnostics, multi-label failure analysis, and a preliminary dynamic-preference split.","tokens_in":19360,"tokens_out":6062,"duration_ms":57880,"significance":"The benchmark addresses a genuine gap by jointly requiring persistent profiles, executable tool-call trajectories, long-horizon interaction, and history-based inference without explicit profile access. The profile-hidden protocol is a solid methodological contribution because it rules out direct profile copying, and the diversity audits for personas and trajectories support internal validity. The paper makes useful falsifiable claims: current frontier tool-use models do not reliably match profile-conditioned reference trajectories, and this inability is not explained by generic tool-use competence. The release of code and prompts is a concrete strength. The central risk is external validity of the LLM-generated reference trajectories; the headline accuracy numbers inherit this risk.","major_comments":[{"comment":"The reference trajectory y* = f_ref(p, h, q, T) is generated by the unspecified 'planner component of the assistant pipeline' (Appendix E), and the human verification interface (Appendix E.1) checks internal consistency, tool validity, and persona consistency rather than agreement with an independently elicited expert decision or with the originating users' preferences. Because exact accuracy is the paper's central metric, the low absolute numbers in Table 4 may partly reflect agreement with this particular generator's planning conventions. The authors should disclose the generator model and version, add a profile-visible oracle condition, and validate a sample of references against independent human judgments with reported inter-annotator agreement.","section":"3.2, Appendix E, Eq. (1)"},{"comment":"The paper properly concedes in Section 3.4 that single-reference exact matching can penalize equally valid preference-compatible trajectories. It then argues from Appendix D that most mismatches are substantive rather than cosmetic. However, the manuscript does not state how the Slight/Major labels in Table 13 were assigned, by whom, or whether annotators agreed, and the multi-label analysis in Table 12 has the same gap. Without annotation protocol and agreement statistics, the diagnostic does not fully separate harmless variation from personalization failure. The authors should publish the labeling protocol, report inter-annotator agreement, or validate a sample of mismatches with independent evaluators.","section":"3.4, Appendix D.2"},{"comment":"All evaluated models are tested only in the profile-hidden condition; there is no profile-visible oracle or upper-bound condition. Such a condition would show whether the reference trajectories are reproducible when the explicit profile is provided and would calibrate the claim that the 49.36% average reflects difficulty of latent preference inference rather than ambiguity or idiosyncrasy of the single reference path. This is a small experiment that is clearly within the scope of the paper and should be added.","section":"4.2"}],"minor_comments":[{"comment":"Several numerical entries in Tables 4 and 5 are typeset without column separators (e.g., '13.4842.22' and '78.5366.80'), making the results difficult to read; the tables should be re-typeset.","section":"4.1, Tables 4-5"},{"comment":"The title on the first page renders 'UserToolBench' as 'USERTOOLBENCH' and 'profile-hidden' as 'HIDDENBENCHMARK' with a missing space; also, Appendix A refers to 'GPT prediction' without specifying which of the evaluated GPT models produced the shown trajectory.","section":"Title, Appendix A"},{"comment":"The multi-label failure percentages in Table 12 are described as jointly assigned, but the manuscript does not define how the labels were derived from failed trajectories; a short description of the labeling procedure would improve reproducibility.","section":"Appendix D.1"},{"comment":"The exact-accuracy definition for lack-of-information tasks credits both clarification and inference, but the evaluation protocol does not state how a partial clarification (e.g., asking for one of several missing constraints) is scored; please clarify this scoring rule.","section":"3.4"}],"recommendation":"major_revision","confidential_remarks":"The core issue for the editor is reference-trajectory validity: the benchmark's central quantitative claims rest on an undisclosed LLM planner as ground truth, and no oracle or external validation is provided. I regard this as fixable within a major revision. The paper fits the journal's scope as a benchmark contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nUserToolBench is worth a serious look. The core idea is simple and new in a way the field needs: build reference trajectories with access to a persistent user profile, then evaluate models without showing them that profile. That hidden-profile split cleanly blocks profile-copying and forces models to actually infer preferences from interaction history. I checked Table 1 and the cited prior work; the combination of executable tool calls, persistent hidden profile, long-horizon interaction, and realistic dialogue is genuinely absent elsewhere. The paper also does a lot of things right: human verification with a described interface, persona diversity audits, a severity taxonomy for failures, and a concrete case showing a model guessing a URL instead of searching history. The result that best exact accuracy is 49% while relaxed accuracy is near 70% is a real finding, and the diagnostics suggest most mismatches are substantive (wrong tools, violated constraints) rather than harmless reorderings. I believe the central claim—that current models complete generic tasks but often miss the personalized decision—holds up in broad strokes.\n\nThe main soft spot is exactly what the stress-test flags: the reference generator is an undisclosed LLM planner. Until we know which model produced the references, the absolute scores and even the exact/relaxed gap could partly reflect agreement with that planner's conventions. The severity audit and the qualitative case soften that worry, but they do not remove it. The paper needs two additions before I'd fully trust the numbers: name the generator, and add a profile-visible oracle baseline. If a model that sees the profile scores near 100% exact, the references are profile-deterministic and the hidden-profile gap is meaningful; if the oracle also lands near 50%, the references may encode generator-specific quirks. The single-reference caveat is acknowledged honestly, and the paper should be credited for that honesty, but it is still a limitation. Next, there are no error bars on the leaderboard and no inter-annotator agreement numbers for the human verification; both are minor and easy to add. The dataset is not yet public, though the authors promise release; that is another minor issue.\n\nThis is a benchmark paper, not a theory paper. Its value is in the protocol and the public resource. The protocol is sound enough to deserve referee time, and the paper is clearly the work of careful people. I would send it to peer review with a request for the oracle and generator disclosure; after that, it is likely a solid contribution. Bring it to the reading group if you care about LLM evaluation or tool-use agents.","headline":"A genuinely new profile-hidden evaluation protocol for personalized tool-use LLMs, with an honest but fixable blind spot around the undisclosed reference generator.","tokens_in":19994,"tokens_out":2762,"would_cite":true,"duration_ms":26889,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"UserToolBench claims that no tested tool-use LLM can reliably recover a profile-conditioned decision trajectory when the explicit user profile is hidden, with the best model scoring 49.36% exact accuracy.","keywords":["personalized decision making","tool-use LLMs","profile-hidden evaluation","preference inference","clarification behavior","tool-call trajectories","multi-tool coordination","long-horizon consistency"],"falsifier":"Ask the users behind the sanitized profiles (or blind judges applying the same preference rules) to rate whether the model predictions that pass relaxed accuracy but fail exact matching are genuinely misaligned with the profile; if a large share are judged equally preference-aligned, the claim that exact-match failures are mostly substantive decision deviations would collapse. A complementary check: regenerate the reference trajectories with a different underlying generator model and measure how much exact accuracy of the same evaluated models changes, since large shifts would show the benchmark tracks one planner's style rather than personalized decision quality.","tokens_in":18948,"feed_emoji":"🎯","tokens_out":7183,"duration_ms":58182,"temperature":0.7,"pith_summary":"This paper introduces UserToolBench, a benchmark that asks whether tool-using large language models can act as personalized delegates: infer a persistent user's latent preferences from interaction history, judge when a request is too incomplete to act on, and emit tool-call trajectories that match what a profile-aware planner would do. The benchmark's signature move is profile-hidden evaluation: reference trajectories are built with the explicit user profile visible and human-verified, while the tested model sees only history, the current request, and tool schemas. Across nine models the best exact-trajectory accuracy is 49.36%, and relaxed task-completion accuracy runs far higher, so models routinely complete the generic task while missing the user-specific decision. The paper's conclusion is that personalization evaluation should measure whether an assistant makes the right executable decision for the represented user, not whether its output sounds user-specific.","feed_headline":"Best tool-use LLM hits 49% when user profiles are hidden","feed_subtitle":"A new benchmark hides user profiles and shows models can finish tasks yet miss the user-specific decision.","key_machinery":"The load-bearing mechanism is the profile-hidden asymmetry: reference trajectories $y^\\star$ are generated as $y^\\star = f_{\\text{ref}}(p, h, q, T)$ with the persistent profile $p$ visible to the generator and validators, while the evaluated model produces $\\hat{y} = f_\\theta(h, q, T)$ with $p$ withheld, forcing preference recovery from interaction history alone. Rounding out the machinery are a milestone-based synthesis pipeline (a persona-conditioned user simulator that deliberately omits decision-critical constraints, a profile-aware planner that fills them from stable preferences, and human verification of tool choice, arguments, clarification behavior, and dependencies) plus a dual scoring scheme (exact trajectory matching and relaxed task-completion accuracy) and failure diagnostics that separate harmless ordering variation from wrong tools, missing calls, violated constraints, and broken dependencies.","core_discovery":"The paper's central claim is that personalized delegation is currently unsolved for tool-use LLMs, and that this failure is visible only when evaluation hides the user profile: with $p$ hidden, the best tested model reproduces the profile-conditioned reference trajectory just 49.36% of the time, while relaxed task-completion accuracy reaches up to 72.55% for the same models. The gap between the two metrics is the paper's key evidence that executable task completion does not imply personalized decision alignment. From failure diagnostics, the authors further argue that exact-match errors are mostly substantive rather than cosmetic, with sequence and dependency errors in 82.1-96.5% of failures, wrong-tool decisions in 48.2-90.3%, and user-constraint violations in 40.9-53.5%, and that multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency are the binding bottlenecks. A preliminary dynamic-preference split shows the same construction framework can represent preference updates, though models still track the latest applicable preference poorly.","pith_inferences":["A direct test of the benchmark's premise would show the same hidden profiles to blind judges and ask whether reference trajectories match stated preferences better than the models' relaxed-compatible alternatives; a high rate of ties would weaken the exact-match interpretation.","The single-reference scoring design suggests an equivalence-class variant: scoring against a set of preference-compatible reference paths, or against a learned preference-satisfaction oracle, would convert the benchmark from agreement-with-one-planner into closer-to-true utility measurement.","The profile-hidden protocol could transfer to other agentic settings such as web navigation, mobile-device control, and code generation, wherever a persistent user's stable constraints must be recovered from history; the paper's own dynamic-preference split points in this direction.","Because the identity of the reference generator is not disclosed, published scores may bound what current models achieve relative to that particular planner's style rather than the ceiling of personalized delegation; disclosing the generator would let readers calibrate the absolute numbers."],"forward_implications":["No tested LLM can yet act as a reliable personalized delegate: exact trajectory accuracy caps at 49.36%, so the capability should be treated as open rather than nearly solved.","Task completion is not a proxy for personalization: because relaxed accuracy runs 20-36 points above exact accuracy on frontier models, benchmarks that score only whether the job got done will systematically overstate personalized alignment.","Multi-tool delegation is the hardest regime (about 25.24% exact accuracy versus 51.30% for single-tool tasks), so personalization difficulty concentrates in sequential decision control, where small early errors propagate through later calls.","Longer interaction history does not by itself improve personalization: several models degrade in later trajectory thirds, indicating that selective retrieval and updating of user state is a distinct capability from accumulating context.","Missing-constraint handling is a separate skill from generic tool use: model rankings on lack-of-information tasks diverge from rankings on single-tool tasks, so benchmarks need dedicated underspecification probes.","A direct test of the benchmark's premise would show the same hidden profiles to blind judges and ask whether reference trajectories match stated preferences better than the models' relaxed-compatible alternatives; a high rate of ties would weaken the exact-match interpretation.","The single-reference scoring design suggests an equivalence-class variant: scoring against a set of preference-compatible reference paths, or against a learned preference-satisfaction oracle, would convert the benchmark from agreement-with-one-planner into closer-to-true utility measurement.","The profile-hidden protocol could transfer to other agentic settings such as web navigation, mobile-device control, and code generation, wherever a persistent user's stable constraints must be recovered from history; the paper's own dynamic-preference split points in this direction."],"supporting_citations":[{"why":"LaMP, the personalization benchmark the paper contrasts with: it evaluates profile-conditioned generation but not executable tool decisions.","marker":"[Salemi et al., 2024]"},{"why":"API-Bank, the tool-use benchmark baseline that grounds the comparison of executable API calling without persistent profiles.","marker":"[Li et al., 2023]"},{"why":"ToolBench (ToolLLM), the large-scale tool-use benchmark whose absence of fixed profiles and long-horizon interaction motivates the new setting.","marker":"[Qin et al., 2024]"},{"why":"Tau-bench, the interactive tool-agent-user environment benchmark that provides the realistic-interaction reference point.","marker":"[Yao et al., 2024]"},{"why":"ToolAlpaca, whose practice of constructing normalized API-style tool schemas is followed when building the benchmark's tool ecosystems.","marker":"[Tang et al., 2023]"},{"why":"PEToolBench, the personalized tool-learning benchmark that covers tool use but not persistent profiles or long horizons.","marker":"[Xu et al., 2025]"},{"why":"PTBench, the personalized tool-invocation benchmark with fixed profiles that still lacks long-horizon reasoning.","marker":"[Huang et al., 2025]"},{"why":"The PersonaMem dynamic-profiling source used to construct the paper's preliminary dynamic-preference split.","marker":"[Jiang et al., 2025]"},{"why":"PersonaLLM, cited for the principle that persona conditions should induce distinct observable behavior, justifying the benchmark's trajectory-diversity measurement.","marker":"[Jiang et al., 2024]"}],"fun_headline_variants":["Tool LLMs pass tasks but fail user intent at 49% match","Hidden profiles expose LLMs' poor personalized tool choice","UserToolBench: 49% match when user profiles are hidden","Tool-use LLMs: task success ≠ personalized decisions","Benchmark reveals gap: 72% task done, 49% user-aligned"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark treats its LLM-generated reference trajectories, built with the profile visible and screened by human verifiers, as the definition of the correct personalized decision, and nothing independently confirms those trajectories match what the real users would actually want.","fun_headline_variants_meta":{"raw":{"variants":["Tool LLMs pass tasks but fail user intent at 49% match","Hidden profiles expose LLMs' poor personalized tool choice","UserToolBench: 49% match when user profiles are hidden","Tool-use LLMs: task success ≠ personalized decisions","Benchmark reveals gap: 72% task done, 49% user-aligned"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000956,"raw_usage":{"total_tokens":4091,"prompt_tokens":976,"completion_tokens":3115,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":3025}},"tokens_in":592,"tokens_out":3115,"duration_ms":21121,"temperature":1.0,"reasoning_tokens":3025,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:13:50.151113+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask the users behind the sanitized profiles (or blind judges applying the same preference rules) to rate whether the model predictions that pass relaxed accuracy but fail exact matching are genuinely misaligned with the profile; if a large share are judged equally preference-aligned, the claim that exact-match failures are mostly substantive decision deviations would collapse. A complementary check: regenerate the reference trajectories with a different underlying generator model and measure how much exact accuracy of the same evaluated models changes, since large shifts would show the benchmark tracks one planner's style rather than personalized decision quality.","supporting_citations":[{"cited_title":"PersonaLLM: Investigating the abil- ity of large language models to express personality traits","cited_arxiv_id":null,"evidence_quote":"PersonaLLM, cited for the principle that persona conditions should induce distinct observable behavior, justifying the benchmark's trajectory-diversity measurement."}],"review_version":1}