{"id":"e5e5ee67-ee02-4ab6-aa9e-e9e068871355","arxiv_id":"2505.03293","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Psi-Arena evaluates LLM counselors through simulated three-phase dialogues with AI clients and tripartite AI judging, then improves them via AI-generated feedback, reporting up to 141% relative pass-rate gains.","lead":"Psi-Arena is a simulated counseling testbed where LLM counselors talk to AI-generated clients and are scored by AI judges playing client, supervisor, and counselor. The authors report model rankings and claim a feedback loop improves counseling pass rates by up to 141%.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 141% improvement is not externally anchored: GPT-4o scores, writes the feedback, and re-scores the same dialogues, while the human validation is rank-order only on 30 dialogues and does not calibrate the absolute pass rates that drive the headline claim.","rationale":"I read the paper as proposing an interactive arena and closed-loop optimization for LLM counselors. For the central claim to hold, the tripartite scores must be a valid measure of counseling quality, and the pre/post improvement must reflect genuine improvement rather than reward hacking against the judge. The weakest point is exactly the self-referential evaluation loop: GPT-4o simulates clients, scores from three perspectives, writes feedback, and re-scores, and it is itself in the evaluated model set. The paper's own validation in Section 3.3 is rank-order-only on 30 dialogues, which cannot calibrate absolute pass rates or effect sizes. My proposed blind clinician study directly tests whether the 141% gain survives an external measure. The reader's weakest_assumption identifies the same load-bearing premise, so I agree with the assessment. The framework is clearly described and plausibly useful as a benchmark, but the headline quantitative claims need this external check; CONDITIONAL remains the appropriate verdict until such validation is available.","tokens_in":21180,"tokens_out":3953,"duration_ms":43866,"concrete_test":"Run a pre-registered blind validation: take the 100 GLM-4-Plus counseling sessions before optimization and the 100 after (or a random sample of 50 pre/post pairs), remove model and condition identifiers, and have three licensed clinicians independently score every dialogue with the paper's three scales in Section 2.4 and Appendix D, applying the same thresholds (>42, >24, >35). Compute overall pass rates under human scoring. The 141% improvement claim is corroborated only if the human-scored relative increase is substantially positive, with a 95% confidence interval excluding zero and overlapping the LLM-judge estimate; if the effect vanishes or reverses, the closed-loop result is an artifact of optimizing against GPT-4o's own grader. Also report inter-rater reliability and a threshold sensitivity sweep of ±10%.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—up to a 141% improvement in counseling performance (Abstract; Section 3.4, GLM-4-Plus overall pass rate 39%→94%)—requires that the tripartite GPT-4o scores faithfully track counseling quality. That premise is load-bearing and is not secured. In Section 2.4 and Appendix C, GPT-4o is used as the simulated client, the three evaluators, and the feedback writer; in Section 2.5 the same model generates the self-reflection feedback that rewrites responses, and GPT-4o is also one of the evaluated counselors in Table 2. The optimizing agent and the measuring instrument are therefore the same function. The only external anchor, Section 3.3, is a rank-order comparison of 30 dialogues from four models by two non-blind psychology graduate students; it validates relative model ordering, not absolute scores, pass thresholds (>42, >24, >35 in Table 1), or pre/post gain sizes. A judge that systematically inflates scores, or consistently prefers its own stylistic preferences, would preserve rank order while changing every pass rate and the 141% figure. Since the pass thresholds are described as \"realistic\" but are not clinically calibrated, and the paper's Limitations section does not address this circularity, the strongest quantitative claims should currently be treated as unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Psi-Arena is an interactive evaluation-and-optimization framework for LLM-based psychological counselors. The authors generate 10,000 GPT-4o-extracted NPC client profiles from the PsyQA dataset, manually select 100 (one per topic), and simulate 25-round, three-phase counseling dialogues between each evaluated LLM counselor and a GPT-4o-simulated client. Each dialogue is scored from three perspectives—client, supervisor, and counselor—by GPT-4o following three rating scales, with pass thresholds (>42/64, >24/32, >35/45) defining competency. Low-scoring dimensions trigger GPT-4o-generated feedback, which the counselor uses to self-reflect and rewrite responses, and revised responses are re-scored by the same pipeline. Eight models are compared (Table 2), a 30-dialogue rank-order human validation is reported (Section 3.3), and the headline finding is that reflection-based optimization yields up to a 141% relative improvement in overall pass rate (GLM-4-Plus, 39% to 94%, Section 3.4). The paper also contributes thematic and dimensional analyses (Section 4).","tokens_in":21613,"tokens_out":10896,"duration_ms":98694,"significance":"The contribution is well scoped and the artifact is genuinely transferable: the paper ships complete prompts (Appendix C), the three rating scales (Appendix D), the topic taxonomy (Appendix B), and a worked case study, so the framework is immediately reusable as a benchmark and as a post-training data generator for mental-health LLMs. The multi-perspective design and the fine-grained thematic and dimensional analyses (Figures 5a-5d) are a clear advance over static knowledge tests and single-perspective client-only evaluations. The 30-dialogue rank-order consistency check is a real, if modest, external anchor. That said, the stress-test concern lands: the evaluation and optimization loop is closed inside GPT-4o (client simulator, three evaluators, feedback writer, and one of the evaluated models), and the human validation calibrates only coarse relative order, not the absolute scores, thresholds, or the 39% to 94% gain that drives the headline. The Limitations section discusses client-simulation complexity, inconsistency of gains, and scalability, but does not acknowledge this circularity.","major_comments":[{"comment":"The evaluation and optimization loop is closed inside GPT-4o. GPT-4o is the client simulator (Section 2.2.3), the evaluator in all three perspectives (Section 2.4; Tables 7-9), the feedback writer (Section 2.5; Tables 10-12), and also one of the evaluated counselors (Table 2). The pre/post comparisons that produce the headline 141% improvement for GLM-4-Plus (39% to 94%) are therefore measurements taken with the same instrument that generated the interventions; a judge that rewards feedback-conformant surface changes (e.g., adding an end-of-session summary, which is item 3 of the counselor scale in Table 17) would inflate gains without any change in clinical quality. The Limitations section does not mention this circularity. The improvement claim needs independent measurement (a separate judge model and/or human absolute scoring on a held-out sample) before it can stand.","section":"Sections 2.4-2.5, 3.4; Appendix C (Tables 7-14); Figure 4"},{"comment":"The human validation is too weak to support the claimed 'high consistency' and does not calibrate the quantities used in the headline. It covers 30 dialogues from four of the eight models; the two raters are non-blind (they are 'very familiar with the research content'), they discuss each dialogue to reach consensus so no inter-rater reliability can be computed, and the analysis reports only rank-order score distributions without any correlation coefficient or confidence interval. This design can at most support a coarse relative-ordering claim; it cannot validate the absolute scores, the three pass thresholds of Table 1, or the gain sizes in Figure 4. The thresholds (>42/64, >24/32, >35/45) are labeled 'realistic' but no clinical or empirical basis is given for them.","section":"Section 3.3; Figure 3; Table 1"},{"comment":"The so-called counselor perspective is not the evaluated model's own reflective awareness. The counselor-perspective scores are produced by GPT-4o role-playing the counselor's self-evaluation (Table 9), not by the evaluated model assessing its own responses. A self-report scale administered to a surrogate is not a measure of the model's reflective practice, so the '360-degree' tripartite framing overstates what is measured. This conflation should be acknowledged explicitly, or the scale relabeled (e.g., 'GPT-4o-predicted counselor self-assessment'), or replaced by the evaluated model's actual self-report.","section":"Section 2.4; Table 9"},{"comment":"The handling of unaddressed dimensions is inconsistent across the three scales and affects the pass-rate metric. The supervisor prompt (Table 8) instructs 'N/A' for dimensions not covered, while the client and counselor prompts (Tables 7 and 9) instruct '0'; in the Appendix F example (Table 21) two supervisor dimensions are scored 'N/A'. The paper does not state whether 'N/A' is mapped to 0 in the totals. This matters because the supervisor threshold is >24 of a nominal 32: with two 'N/A' dimensions the achievable maximum is 24, so the threshold means something different from what Table 1 implies. Since overall pass rate requires meeting all three thresholds, this ambiguity propagates directly into Tables 2 and Figure 4.","section":"Table 1; Tables 7-9; Appendix F (Table 21)"},{"comment":"The abstract's claim of 'significant performance variations' across models is not supported by any significance test, confidence interval, or multiple-comparison control; the per-dimension scores in Figure 5 span 0-5 for all models, so dialog-level variance is clearly large. Likewise, the pre/post improvements in Figure 4 are presented without per-client variance, paired tests, or bootstrap intervals, so the reader cannot determine which reported gains exceed the noise of the GPT-4o judge.","section":"Section 3.1; Table 2; Figure 4"}],"minor_comments":[{"comment":"The framework name is inconsistent: the title and most of the abstract use 'Psi-Arena,' the final sentence of the abstract says 'PsychoArena,' and the section headings are typeset as 'ARENA'; please standardize.","section":"Abstract and headings"},{"comment":"The text says the counselor scale 'covers 20 dimensions, with 9 focused on practical counseling abilities,' but Table 1 and Table 17 present a 9-item scale; please clarify whether the administered instrument had 9 or 20 items and where the other 11 dimensions are described.","section":"Section 2.4; Table 1; Table 17"},{"comment":"The axis labels and caption of Figure 3 are not legible in the provided manuscript, and the comparison would be far more informative as a quantitative rank-correlation (e.g., Spearman's rho with a confidence interval) reported per scale.","section":"Figure 3"},{"comment":"The manuscript reports generating 10,000 profiles but using one manually selected profile per topic (100 clients for evaluation); please describe how the single profile per topic was chosen and acknowledge the potential selection bias this introduces into every downstream result.","section":"Section 2.2.1"},{"comment":"Since all counselors interact with the same GPT-4o client simulator, measured differences across counselors confound counselor ability with the simulator's reactions to different conversational partners; a sentence acknowledging this interaction confound would be helpful.","section":"Section 2.2.3 and Section 3.1"},{"comment":"The three rating scales are imported from prior sources (Black 2003; APA 2023; Yang and Xiong 2018) without reporting their psychometric properties in the adapted English form; please cite the original validation studies and note any translation or adaptation that was performed.","section":"Section 2.4 and Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for the journal and the framework artifact is genuinely reusable, but the headline quantitative claims (the 141% improvement and the pass-rate rankings) are currently measured by the same GPT-4o system that generates the interventions, and the 30-dialogue human validation is rank-order only. I recommend major revision with a requirement to re-anchor the evaluation: absolute-score human calibration on a held-out sample, a quantitative agreement metric with confidence intervals, an independent judge (human or a second model) for the pre/post comparison, and clarification of the N/A scoring rule. The abstract's naming inconsistency and unqualified 'significant' claims suggest the manuscript would also benefit from a careful pass before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Yue, quick read of Psi-Arena. The thing to know: this is a genuinely useful assembly job, not a breakthrough. The tripartite evaluation (client, supervisor, counselor perspectives) over multi-stage simulated dialogues is a real step beyond the static knowledge tests and single-perspective evaluations that dominate this corner of the literature. The 100 client profiles, the three-phase interaction design, the 33 evaluation dimensions, and the closed-loop feedback mechanism are all clearly described and well documented in the appendix. The authors also deserve credit for shipping the full prompt suite and a dialogue sample. The framework is reusable.\n\nThe soft spot is exactly where the stress-test note lands. The optimization loop is closed inside GPT-4o: the same model plays the client, the three evaluators, the feedback writer, and in one condition the evaluated counselor. The human validation is rank-order only, on 30 dialogues from four models, by two non-blind graduate students who also helped build the evaluation criteria. That anchors relative ordering, not absolute scores, not the pass thresholds, and not the pre/post gain sizes. So the \"141% improvement\" (GLM-4-Plus: 39% to 94%) should be read as a self-consistent simulation result, not as measured counseling improvement. The Limitations section is candid about client realism and scalability but never mentions this circularity, which is a real omission.\n\nThat said, the central architectural idea holds up. The framework is not pretending to be a clinical trial; it is a simulated arena for model comparison and iterative improvement. For that purpose, the paper is solid and the documentation is unusually thorough. The rank-order human check is a reasonable first step, though inter-rater reliability and a blind setup would make it much stronger. I would not treat the absolute pass rates as clinically meaningful, and I would not cite the 141% number as evidence about counseling quality. I would cite the framework itself as a useful evaluation harness and would absolutely send this to peer review. A serious referee should push for artifact release, a blinded and larger human validation, and a clearer separation between the judge model and the evaluated models in the optimization loop.","headline":"A well-built tripartite evaluation harness for LLM counselors that deserves serious peer review, but the headline 141% improvement is not yet anchored by external validation.","tokens_in":22071,"tokens_out":531,"would_cite":true,"duration_ms":6887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Psi-Arena claims that tripartite feedback lifts LLM counseling pass rates by up to 141 percent.","keywords":["LLM evaluation","psychological counseling","tripartite feedback","NPC client simulation","closed-loop optimization","mental health","self-reflection"],"falsifier":"Have an independent panel of licensed counselors, blind to model identity and to whether a response is pre- or post-feedback, score the same sessions on the three published scales; if their absolute scores and pass rates do not reproduce the reported ranking and the 39% to 94% improvement, the claim that the loop improves counseling competence would be falsified.","tokens_in":20952,"feed_emoji":"🧠","tokens_out":7077,"duration_ms":61561,"temperature":0.7,"pith_summary":"Psi-Arena's central claim is that LLM counseling competence should be evaluated in dynamic, multi-turn sessions rather than static tests, and judged from three perspectives at once. The paper builds a closed assessment environment in which psychologically profiled simulated non-player-character (NPC) clients hold 25-round counseling dialogues with eight LLM counselors, and GPT-4o plays client, supervisor, and counselor to score each session on established scales. It then argues that low-scoring dimensions can be turned into concrete feedback, and that self-reflection on that feedback makes counselors measurably better. The headline result is a 141% relative improvement in overall pass rate for one model, with weaker models improving most. If correct, this gives mental-health LLM developers a reusable loop for benchmarking and improving counseling ability rather than only knowledge retention.","feed_headline":"Tripartite feedback lifts LLM counselor pass rates by up to 141%","feed_subtitle":"A simulated client, supervisor, and counselor grade each session; self-reflection turns weak replies into stronger ones.","key_machinery":"The load-bearing mechanism is the tripartite evaluation loop. Each 25-round dialogue is scored by GPT-4o playing three roles: a simulated client using a 16-item scale (0-4, pass threshold >42/64), a supervisor using an 8-item competence scale (0-4, threshold >24/32), and the counselor itself using a 9-item self-assessment scale (0-5, threshold >35/45). Low-scoring dimensions are converted into feedback grounded in 11 counseling approaches, and the counselor reflects on that feedback and rewrites its response before being re-scored. The framework's realism comes from NPC clients whose profiles are extracted from real counseling records and whose behavior follows trust-building, diagnostic, and solution-exploration phases. These scales and thresholds jointly define the pass rate that drives the paper's improvement claims.","core_discovery":"The paper's central discovery is that a tripartite, feedback-driven arena can reveal and improve counseling skills that single-perspective, static evaluations hide. Using 10,000 NPC client profiles built from real counseling records, it stages multi-stage dialogues and scores each one from the client (subjective experience), supervisor (professional competence and ethics), and counselor (reflective self-awareness) perspectives across 33 dimensions. The authors report that the three perspectives disagree in informative ways, that the resulting diagnostic feedback produces consistent gains across all eight models, and that GLM-4-Plus's overall pass rate rises from 39% to 94% (a 141% relative improvement). Automated rankings also track the rankings of two psychological experts on a 30-dialogue sample. The claim is that this makes Psi-Arena a valid and reusable testbed for LLM-based psychological counseling.","pith_inferences":["Editorial: Because the judge and the client are both GPT-4o, the reported gains could partly reflect the counselor learning to satisfy GPT-4o's preferences; a blind human-scored pre/post comparison would separate genuine counseling gains from judge adaptation.","Editorial: The tripartite scheme could be transplanted to non-Chinese cultural contexts, crisis intervention, or diagnostic accuracy tasks, but the thresholds and scales would need recalibration for those settings.","Editorial: The fixed three-phase, 25-round structure may reward formulaic stage compliance—such as ending with a summary—rather than individualized responsiveness; testing with varied client profiles and session lengths would reveal this.","Editorial: Counselor self-assessment in the loop may encourage models to narrate competence rather than demonstrate it; coupling the self-score with observable behavioral outcomes would make the loop harder to game."],"forward_implications":["Static knowledge tests and single-perspective user satisfaction measures understate differences between counselors; multi-turn tripartite evaluation reveals which models actually sustain a therapeutic relationship.","Supervisor-style criteria are the strictest gate, so models that pass only client-satisfaction checks may still fail on professional competence and ethics.","Closed-source and open-source models vary widely by topic and dimension, with treatment- and career-related themes being the hardest for most models.","Feedback-driven self-reflection improves all tested models and helps weaker models more, so the loop offers a practical way to upgrade counseling LLMs without retraining.","Stronger models benefit more from a simple psycho prompt, suggesting that instruction-following capacity bounds how much role framing can improve counseling."],"supporting_citations":[{"why":"Provides the real-world PsyQA counseling records from which the 10,000 NPC client profiles are extracted.","marker":"Sun et al., 2021"},{"why":"Supplies the 16-dimension client evaluation scale used for the client perspective.","marker":"Black, 2003"},{"why":"Supplies the 8-dimension supervisor competency assessment tool used for the supervisor perspective.","marker":"APA, 2023"},{"why":"Supplies the counselor self-assessment scale whose nine ability dimensions form the counselor perspective.","marker":"Yang and Xiong, 2018"},{"why":"Supplies the 11 counseling-theory knowledge base used to ground the generated feedback in professional approaches.","marker":"Corey, 2013"},{"why":"Defines the solution-exploration stage that virtual clients are instructed to follow in the final phase of the dialogue.","marker":"Hill, 2020"}],"fun_headline_variants":["Arena feedback lifts LLM counselor pass rates by 141%","Client, supervisor, and counselor feedback boosts LLM therapy skills","Interactive arena reveals gains in LLM counseling, up to 141%","Tripartite feedback drives LLM counselor pass rates from 39% to 94%","LLM counselor pass rate jumps 141% with tripartite feedback"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that GPT-4o, when prompted as the client, supervisor, and counselor, scores dialogues the way trained human judges would, even though only a small rank-order comparison against two expert judges is used to check this.","fun_headline_variants_meta":{"raw":{"variants":["Arena feedback lifts LLM counselor pass rates by 141%","Client, supervisor, and counselor feedback boosts LLM therapy skills","Interactive arena reveals gains in LLM counseling, up to 141%","Tripartite feedback drives LLM counselor pass rates from 39% to 94%","LLM counselor pass rate jumps 141% with tripartite feedback"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000901,"raw_usage":{"total_tokens":3868,"prompt_tokens":921,"completion_tokens":2947,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":2850}},"tokens_in":537,"tokens_out":2947,"duration_ms":20088,"temperature":1.0,"reasoning_tokens":2850,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:55:00.218462+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent panel of licensed counselors, blind to model identity and to whether a response is pre- or post-feedback, score the same sessions on the three published scales; if their absolute scores and pass rates do not reproduce the reported ranking and the 39% to 94% improvement, the claim that the loop improves counseling competence would be falsified.","supporting_citations":[],"review_version":1}