{"id":"92613fc2-e6d0-46f0-9d12-6105a62892e7","arxiv_id":"2505.18953","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":8,"one_line_summary":"An evaluation of eight LLMs on a hand-built risk-tolerance rubric finds large scoring errors and small demographic differences, but the rubric given to the models contradicts the rubric used as ground truth.","lead":"This paper asks whether eight AI chat models can score an investor's risk tolerance from a written profile, and whether the scores shift when the person's country or gender changes. The authors report that no model is both accurate and consistent across demographics, but the benchmark's own scoring rules are internally inconsistent, weakening the conclusions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Country-level bias findings may be confounded: the benchmark does not establish that non-demographic financial features are held constant across countries, so score gaps could reflect legitimate financial differences rather than demographic bias.","rationale":"The paper's most consequential empirical claim is not merely that scores deviate from a rubric; it is that models are biased by country and gender, violating finance regulations. That claim rests on the consistency analysis in Tables 3–4. For those comparisons to isolate protected attributes, all 16 financial features must be held fixed across countries and genders. The manuscript does not say they are, and Section 4.1's country grouping by economic characteristics suggests profiles carry country-correlated financial information. Without matched profiles or adjustment, the observed country gaps (e.g., Nigeria/Indonesia higher) are confounded.\n\nThe rubric mismatch between Section 4 and Appendix A.6 identified by the reader is real and invalidates the correctness axis, but the consistency claim could in principle survive a corrected rubric; profile confounding, if present, destroys the bias claim directly. The reader's broader rationale does mention that profile construction is underspecified, so I partially agree with the reader's assessment. A concrete matching/adjustment check would settle the point.\n\nEither way, the paper as it stands does not support the regulatory-violation conclusion, so the REJECT verdict is appropriate and my concern does not change it.","tokens_in":13802,"tokens_out":5366,"duration_ms":53208,"concrete_test":"Reconstruct or release the profile-generation code and extract the 16 financial features for every profile. Test balance across country and gender (e.g., standardized mean differences or ANOVA). If features are imbalanced, re-run Tables 3 and 4 using regression with the 16 features as covariates, or on a matched subset. If adjusted country/gender differences drop below the reported 0.3–1.0 point range and lose significance, the bias claim fails. If the profiles are already balanced, this test would confirm the concern is not valid.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central consistency claim ('none maintain consistent scores across regions and demographics') depends on treating country and gender as isolated variables. The paper does not establish that profiles are matched on the 16 financial features. Section 4.1 says only that 'for each of the 10 selected countries, we selected two representative names per gender' and 'for each name, we constructed 43 summaries'; it never states that income, assets, expenses, debt, investment tenure, etc. are identical across countries/genders. The same section classifies countries into groups 'based on population size and economic characteristics', noting that individuals in populous countries often have 'limited financial history', while those in less populous countries have 'strong banking infrastructure'. If these financial-history differences appear in the interview_text, then higher scores for Nigeria/Indonesia in Table 3 can be a correct response to lower financial capacity, not bias. No matching, covariate adjustment, or ablation is reported. Hence the headline finding of demographic bias, and the regulatory-violation conclusion, is not supported by the data as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FINRISK EVAL, a benchmark of 1,720 synthetic user profiles spanning 10 countries and two genders, each described by 16 financial features plus demographic attributes. A hand-built scoring rubric (Section 3 and 4) assigns each feature a numeric value and sums them into a Risk Tolerance (RT) score, with profiles categorized as Conservative, Moderate, or Aggressive. Eight LLMs (GPT-4o, GPT-4o mini, Gemini 1.5 Pro, Claude 3.7 Sonnet, DeepSeek-V3, LLaMA 3.1 405B, LLaMA 3.3 70B, and Mistral Small) are prompted to output an RT score for each profile. The paper then analyzes correctness (deviation from scenario-level ideal scores) and consistency (variation in predicted scores across countries and genders). It reports that while some models align with ideal scores in low- and mid-risk scenarios, none is both accurate and consistent across all demographic groups, and it concludes that current LLMs are not ready for automated risk profiling and that their inconsistencies violate AI and finance regulations.","tokens_in":13984,"tokens_out":9813,"duration_ms":80956,"significance":"If the evaluation were valid, the paper would provide a valuable public benchmark for auditing LLMs in a high-stakes financial task. The dataset construction (1,720 profiles, 10 countries, balanced gender) and the explicit separation of demographic from financial features are useful ideas, and the authors test a diverse set of proprietary and open-weight models. The paper also makes a concrete, falsifiable prediction: no current LLM produces risk scores that are simultaneously accurate and stable across demographics. However, the evaluation has several load-bearing methodological flaws that invalidate the reported numbers: the prompt rubric differs from the ground-truth rubric on a key feature, the correctness metric compares against arbitrary scenario-level ideals rather than per-profile ground truths, and the consistency analysis does not control for potentially different financial features across countries and genders. The circularity between the ground-truth rubric and the prompt further weakens the claim that the study assesses models' independent risk-assessment ability. As a result, the paper's central empirical claims are not supported.","major_comments":[{"comment":"The ground-truth scoring rubric and the prompt rubric disagree on the 'Investment Amount Monthly' feature. Section 4 states that 'contributing less than 10% of income monthly suggests lower risk tolerance (+2 points), whereas investing more than 30% implies higher exposure and results in -2.' Appendix A.6 instructs models to award '+1 (>20%), 0, -1 (<20%)' for the same feature. A model that follows the prompt will therefore produce scores that cannot match the ground-truth values, and the deviation scores in Tables 1 and 2 do not measure the model's accuracy. The authors must either use identical rubrics or analyze the discrepancy explicitly.","section":"Section 4 vs. Appendix A.6"},{"comment":"The correctness analysis compares each country's mean predicted score to fixed 'ideal' values (-5, 10, 21.5) for the Low, Mid, and High scenarios, rather than to the profile-specific ground-truth RT scores defined in Section 3. Because each profile has its own computed RT score, the deviation should be computed as predicted score minus ground-truth score per profile and then aggregated. The chosen ideals are not derived from the dataset (they are not the midpoints of the stated ranges), and they ignore the distribution of ground-truth scores within each scenario. Consequently, the reported 'mean difference' values do not indicate how well a model predicts individual risk tolerance.","section":"Section 6.1, Tables 1 and 2"},{"comment":"The consistency analysis treats country and gender as isolated variables, but the paper never states that profiles are matched on the 16 financial features across these groups. Section 4.1 distinguishes 'highly populous' from 'less populous' countries and notes that individuals in the former 'lack access to formal banking systems' and have 'limited financial history,' implying the interview_text may contain different financial information. If so, higher scores for Nigeria and Indonesia (Table 3) could reflect legitimate differences in income, assets, or expenses rather than demographic bias. No matching, covariate adjustment, or ablation is reported, so the country- and gender-bias claims are not supported.","section":"Sections 4.1 and 6.2-6.4"},{"comment":"The evaluation is circular with respect to the paper's central research question. The ground-truth RT score is defined by the same hand-built rubric that is placed verbatim into the model prompt, so a model that adheres to the prompt will reproduce the authors' scoring formula. The resulting 'correctness' scores therefore measure instruction-following, not the model's independent ability to assess investment risk appetite. The paper's framing ('Can AI Credible at Assessing Investment Risk?' and the claim that models fail to 'accurately predict an individual's risk appetite') is not supported by this design.","section":"Introduction, Section 3, and Appendix A.6"},{"comment":"The claim that the observed inconsistencies 'violate AI and finance regulations' is not substantiated. The paper does not identify a specific regulatory provision or standard that the models' score differences violate, nor does it report any statistical test (e.g., significance of country/gender differences) that would support a finding of discrimination. Differences of 0.3-1.0 points, as reported in Sections 6.3-6.4, may be within the noise of the scoring procedure; without inferential statistics or a regulatory benchmark, the violation claim is overreaching.","section":"Abstract and Section 7"}],"minor_comments":[{"comment":"The rubric and prompt also differ on several features beyond the monthly investment amount (e.g., Age is given as '+2, +1, 0' in the prompt without the thresholds specified in Section 4), and the prompt's allowed range of -15 to 28 is inconsistent with the -14 to 28 range defined in Section 4.","section":"Section 4 vs. Appendix A.6"},{"comment":"The list of evaluated models is confusing: Section 5 says 'LLaMA 3.1 (70B and 405B)' were evaluated, but later in the same section it refers to 'LLaMA 3.1 (8B/70B)' among skipped models; please clarify which LLaMA variants were actually used in the reported results.","section":"Section 5"},{"comment":"The paper would benefit from including at least one example of the interview_text for a profile, so readers can verify that all 16 features are represented and that country and gender are the only varying attributes across the constructed profiles.","section":"General"},{"comment":"Table 4 is very dense and hard to parse; a separate column showing the female-male difference per country, or smaller tables per scenario, would help the reader assess the reported gender effects.","section":"Table 4"},{"comment":"The text describes mean differences such as 7.27 as 'deviation' without specifying whether these are signed or absolute deviations; since all Low-scenario entries are positive in Table 1, the sign indicates overestimation relative to the ideal, but the caption and text should state this explicitly.","section":"Section 6.1"}],"recommendation":"reject","confidential_remarks":"The paper is better suited to a benchmark/evaluation venue, but even there the methodology needs major revision. In particular, the authors should re-run experiments with a consistent rubric and report profile-level accuracy against ground-truth scores, and they should either match profiles on financial features or use regression adjustment for the consistency analysis. If these changes are made, the benchmark could be useful."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nPunchline: the paper asks a timely question about whether LLMs can assess investment risk appetite, but the benchmark it builds doesn't support the headline claim. The correctness metric is invalidated by a mismatch between the ground-truth rubric and the prompt rubric, and the country-level consistency results are likely confounded because profiles aren't matched on financial features.\n\nWhat's good: it addresses a real deployment problem with a wide range of proprietary and open-weight models, and the structured rubric is grounded in regulatory frameworks. The decision to explicitly include demographic attributes to test whether models rely on them is sensible. The paper also reports detailed country and gender breakdowns, and the authors acknowledge limitations in the appendix.\n\nThe soft spots are serious. In Section 4, the ground truth for 'monthly investment amount' gives +2 for contributing less than 10% of income and -2 for more than 30%; in Appendix A.6, the prompt tells models to give +1 for more than 20% and -1 for less than 20%. So models are graded against rules they weren't fully shown. This alone undermines the deviation scores. More broadly, 'correctness' is defined relative to a hand-built rubric that is also placed in the prompt, so the evaluation largely measures instruction following, not real-world accuracy. The consistency analysis doesn't establish that the 16 financial features are held constant across countries; the paper itself groups countries by population and economic characteristics, so higher scores for Nigeria/Indonesia could reflect legitimate lower financial capacity rather than bias. There are no significance tests, no repeated sampling, and the paper drops models that didn't follow instructions. The dataset isn't released.\n\nThe abstract overstates the results: the best low-risk deviation is about 7 points, and gender differences are sub-point, so claiming 'violating AI and finance regulations' is too strong.\n\nWho's this for? Researchers studying LLM fairness in finance could use the benchmark as a starting point, but only after fixing the rubric inconsistency and adding matched profiles. As is, the central claim isn't supported. I'd recommend rejecting this version, but with a clear path to revision: align the prompt and ground-truth rubrics, match or covary the financial features, and add uncertainty quantification. It deserves a serious referee only after those fixes.","headline":"A timely question undermined by a rubric mismatch and unmatched profiles; the benchmark doesn't support the claim that no LLM is credible.","tokens_in":14603,"tokens_out":3264,"would_cite":false,"duration_ms":27989,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No AI system consistently produces unbiased, accurate investment-risk scores across all demographic groups and countries.","keywords":["LLM evaluation","investment risk appetite","demographic bias","fairness","financial regulation","risk tolerance scoring","consistency"],"falsifier":"Recompute the ground-truth scores for all 1,720 profiles using the rubric actually printed in Appendix A.6 rather than the one in Section 4; if the per-model deviation and consistency tables change materially, the reported correctness and bias findings depend on the scoring mismatch. A second decisive check is a token-swap experiment: take one profile and change only the country and name tokens while leaving the other sixteen features untouched; the paper's claim predicts systematic score shifts across every model, and their absence would refute it.","tokens_in":13523,"feed_emoji":"⚖️","tokens_out":8376,"duration_ms":60063,"temperature":0.7,"pith_summary":"This paper asks whether current large language models can credibly score how much investment risk a person should take on, and its answer is no. On a 1,720-profile benchmark whose ground-truth scores come from an additive five-part rubric, all eight tested models shift their scores when a user's country or gender changes, even though those attributes are supposed to be irrelevant. The paper defines credibility as correctness (closeness to the rubric's ideal score) plus consistency (stability across demographic groups), and it reports that no model achieves both. GPT-4o tracks low- and mid-risk targets closely but scores Nigerian and Indonesian profiles higher; open-weight models show irregular gender gaps. If the results hold, current LLMs are not ready for automated investment risk profiling without demographic fairness controls.","feed_headline":"No AI model passes the investment-risk fairness test","feed_subtitle":"A 1,720-profile benchmark shows every tested LLM shifts scores with country and gender.","key_machinery":"The load-bearing object is FINRISK EVAL, a benchmark of 1,720 user profiles whose ground-truth risk score is computed by an additive rubric: RT = PFS + ISO + LAA + MCR + DOI, one term for each of five dimensions: personal and financial stability, investment strategy and objectives, liquidity and asset allocation, market and currency risks, and dependency on investments. Each profile pairs fixed financial features with a country and a binary gender, so any difference in model output between otherwise matching profiles is attributable to a demographic attribute the rubric itself ignores. This design turns demographic bias into a measurable quantity: the cross-country standard deviation of the model's mean score and the gender gap in each scenario.","core_discovery":"On the paper's own terms, the central discovery is that none of the evaluated AI systems—GPT-4o, GPT-4o mini, Gemini 1.5 Pro, Claude 3.7 Sonnet, LLaMA 3.1 (405B), LLaMA 3.3 (70B), DeepSeek-V3, and Mistral small—consistently produces unbiased, accurate risk-tolerance scores across all demographic groups and countries. Each model's score is compared against the rubric-based ideal for conservative (low), moderate (mid), and aggressive (high) profiles; the deviations and cross-country standard deviations are the evidence. The authors find certain models align well in specific ranges, but no model keeps its output stable when country or gender changes, and the authors read this instability as a failure of financial fairness and regulatory compliance. The concrete examples include GPT-4o assigning higher risk scores to Nigerian and Indonesian profiles and open-weight models exhibiting inconsistent gender-based scoring.","pith_inferences":["Beyond the paper, the rubric in Section 4 and the prompt in Appendix A.6 disagree on at least the monthly-investment rule, so the reported deviation numbers may grade models against a standard they were not asked to follow; rerunning with a single consistent rubric would separate true miscalibration from prompt artifacts.","Without a human-advisor baseline, the same benchmark could also be read as evidence that risk-appetite scoring is inherently noisy; comparing human advisors on identical profiles would clarify whether the observed spread is an AI-specific failure.","The profile design swaps names and countries embedded in the interview text, so the test measures sensitivity to the whole demographic token set; a fully controlled experiment would vary just one token at a time to estimate causal effects of each attribute.","The benchmark could be extended to non-binary gender identities, more countries, and other protected attributes, and could be used to test whether fairness fine-tuning actually shrinks the reported inconsistencies."],"forward_implications":["Regulators should treat demographic consistency as a precondition for deploying LLMs in suitability and risk-profiling contexts.","Model choice should be risk-band specific, since alignment on low and mid profiles, as with GPT-4o, does not carry over to aggressive profiles.","Fine-tuning and calibration targets should include demographic invariance rather than only average correctness.","Deploying these models without such controls creates legal exposure under frameworks like the EU AI Act, GDPR, the Equal Credit Opportunity Act, and MAS Fairness, Ethics, Accountability and Transparency principles.","A standardized public benchmark built on this profile structure could let firms audit future models the same way."],"supporting_citations":[{"why":"Supplies the regulatory suitability features—income, assets, liabilities, and objectives—that define the profile dimensions.","marker":"FINRA, 2012"},{"why":"Provides the regulatory basis for considering financial position and risk preferences in suitability assessment.","marker":"FCA, 2018"},{"why":"Establishes the suitability requirement that financial situation and investment goals be accounted for, shaping the rubric.","marker":"ESMA, 2023"},{"why":"Sets the Singapore regulatory expectation that risk profiling use income, assets, liabilities, and objectives, and grounds the fairness standards cited.","marker":"MAS, 2023"},{"why":"Supplies the risk-tolerance instrument whose variables the benchmark's feature set is adapted from.","marker":"Grable and Lytton, 1999"},{"why":"Justifies portfolio diversification as a risk factor in the rubric.","marker":"Markowitz, 1952"},{"why":"Prior evidence that LLM investment advice carries risk biases, motivating the correctness axis.","marker":"Guo et al., 2024"},{"why":"Prior evidence of systematic product bias in LLM investment recommendations, motivating the consistency axis.","marker":"Zhi et al., 2025"}],"fun_headline_variants":["Every AI model fails investment-risk fairness test","AI risk scores shift with country and gender in all models","No LLM keeps risk scores stable across demographics","All tested AI systems flunk unbiased risk assessment"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that the scoring rubric used to compute ground-truth risk scores is the same rubric the models are given in the prompt, but the paper contradicts this for the monthly-investment feature: the scoring section awards +2 for investing under 10% of income and −2 for over 30%, while the prompt awards +1 for over 20% and −1 for under 20%.","fun_headline_variants_meta":{"raw":{"variants":["Every AI model fails investment-risk fairness test","AI risk scores shift with country and gender in all models","No LLM keeps risk scores stable across demographics","All tested AI systems flunk unbiased risk assessment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000104,"raw_usage":{"total_tokens":988,"prompt_tokens":858,"completion_tokens":130,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":70}},"tokens_in":474,"tokens_out":130,"duration_ms":1812,"temperature":1.0,"reasoning_tokens":70,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:23:44.152716+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the ground-truth scores for all 1,720 profiles using the rubric actually printed in Appendix A.6 rather than the one in Section 4; if the per-model deviation and consistency tables change materially, the reported correctness and bias findings depend on the scoring mismatch. A second decisive check is a token-swap experiment: take one profile and change only the country and name tokens while leaving the other sixteen features untouched; the paper's claim predicts systematic score shifts across every model, and their absence would refute it.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the regulatory suitability features—income, assets, liabilities, and objectives—that define the profile dimensions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the regulatory basis for considering financial position and risk preferences in suitability assessment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the suitability requirement that financial situation and investment goals be accounted for, shaping the rubric."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets the Singapore regulatory expectation that risk profiling use income, assets, liabilities, and objectives, and grounds the fairness standards cited."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Justifies portfolio diversification as a risk factor in the rubric."},{"cited_title":"Nucleation and growth manifest universal scaling, surely","cited_arxiv_id":"2405.11231","evidence_quote":"Prior evidence that LLM investment advice carries risk biases, motivating the correctness axis."}],"review_version":1}