REVIEW 4 major objections 5 minor 99 references
Understanding and Mitigating Risks of Generative AI in Financial Services
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read General-purpose AI guardrails fail to detect most content risks specific to financial services, and prompting them with a new 12-category risk taxonomy does not close the gap.
desk verdict A useful financial-services risk taxonomy with a plausible but under-supported empirical claim; the missing control set weakens the central 'safety gap' conclusion. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the paper's financial-services AI content risk taxonomy: twelve categories (Confidential Disclosure, Counterfactual Narrative, Defamation, Discrimination, Financial Services Impartiality, Financial Services Misconduct, Irrelevance, Non-Financial Advice, Offensive Language, Personally Identifiable Information, Prompt Injection and Jailbreaking, Social Media Headline Risk) plus jurisdiction- and product-specific overlays, grounded in the obligations of buy-side firms, sell-side firms, and technology vendors. The taxonomy carries the argument because it converts regulatory and fiduciary duties into a concrete definition of unsafe content, and the evaluation protocol—red-teamed inputs and outputs annotated by majority vote, guardrails run in default and prompt-expanded settings, and scored by recall, precision, and F1 against the taxonomy—turns that definition into a quantifiable benchmark that exposes the safety gap.
What would settle it
If a re-annotation of the red-teaming dataset with the full annotation guide, a third adjudicator, and a measured agreement score moves the best guardrail's recall on Financial Services Impartiality or PII from below 0.4 to above 0.8, the reported safety gap would be an artifact of label noise.
Extended reading notes
Core claim
The central discovery is an empirical safety gap: existing LLM-based guardrails fail to detect most of the content risks that matter in financial services. Evaluated on 10,400 red-teamed system inputs and 7,340 system outputs against the paper's financial-services taxonomy, all four guardrails achieve high precision but low recall on inputs, and perform poorly on outputs in both precision and recall. Prompting the models with the new risk categories does not overcome the limitation; ShieldGemma's expanded version raises recall but drops precision and drives its false-positive rate on normal business queries to 32.8 percent. The paper attributes the gap to the fact that safeguards are fine-tuned to recognize their own general-purpose taxonomies and were not designed to cover other definitions of risk natively or through prompts, concluding that general taxonomies are a starting point but not sufficient for domain-specific GenAI systems.
Load-bearing premise
The evaluation's ground-truth labels are accurate enough to support the recall and F1 numbers, even though 2,337 of the 17,740 examples were labeled by only two experts with no detailed annotation instructions and no inter-annotator agreement was reported.
Editorial extensions
If this is right
- Deploying general-purpose guardrails without domain adaptation on financial GenAI systems will leave most domain content risks undetected, so system builders need domain-specific guardrails or additional mitigation layers.
- Prompt expansion alone is not a workable adaptation strategy: adding the new risk categories to the guardrail prompts failed to improve recall consistently and, for two models, raised the false-positive rate on normal business queries from near zero to 5.2 and 32.8 percent.
- In-domain performance claims are misleading for domain transfer: models that report F1 scores of 0.76–0.94 on their own test sets drop to 0.34–0.58 on the overlapping 'social media headline risk' category when evaluated on financial-services examples.
- Risk categories in a domain taxonomy must be precise and grounded in the applicable legal and regulatory context, because categories that share a name across taxonomies (e.g., discrimination, PII) can differ enough to change guardrail behavior.
Reading between the lines
- Beyond the paper's claims, the same safety gap plausibly extends to other regulated knowledge-intensive domains such as healthcare and law, where general-purpose guardrails are deployed without domain adaptation; the paper's method of deriving risk categories from stakeholder duties is portable.
- The false-positive results on normal business queries suggest a deployability consequence not highlighted by the paper: in settings where malicious queries are rare, even a modest false-positive rate means the guardrail blocks far more legitimate queries than it catches, making the measured trade-off a cost problem as much as a safety problem.
- A testable extension, not pursued by the paper, is whether a small classifier fine-tuned on the taxonomy's examples outperforms all four LLM guardrails at lower latency and cost, which would re-frame the gap as a training-data problem rather than an inherent limitation of LLM-based moderation.
- If the red-team data were released as a static benchmark, the paper's one-time evaluation would become a regression test that future domain-adapted guardrails must pass, a natural follow-up implied by the paper's recommendations.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that general-purpose AI safety taxonomies and guardrail systems are insufficient for regulated, knowledge-intensive domains, and it develops a financial-services-specific content risk taxonomy grounded in the duties of buy-side firms, sell-side firms, and technology vendors. It then evaluates four open-source guardrail systems (Llama Guard, Llama Guard 3, AEGIS, and ShieldGemma) on red-team inputs and outputs labeled with this taxonomy, reporting high precision but low recall in default and prompt-expanded configurations, and concludes that a 'safety gap' arises from the guardrails' inability to recognize financial-services risks.
Significance. If the empirical conclusion is supported, the paper makes a useful contribution by shifting content-safety evaluation from model-level, general-purpose taxonomies toward domain-specific sociotechnical risk assessment. The taxonomy itself is careful and well grounded in legal and regulatory source material, and the paper explicitly contrasts its stakeholder-based approach with existing taxonomies such as MLCommons. The evaluation also has a practical strength: it uses a large red-teaming dataset (10,400 inputs and 7,340 outputs) and reports precision, recall, and false-positive rates rather than only benchmark accuracy. The paper's recommendations on governance, multi-layer safety, and domain-adapted guardrails are reasonable. However, the causal claim that the observed low recall is specifically a financial-services domain gap is not yet established by the evidence presented.
major comments (4)
- [§6, Table 2 and Table 3] The central claim that the safety gap 'emerges from the guardrails' inability to recognize many sources of risk in the financial services domain' is not supported without a hard-positive general-domain control. The red-team dataset contains only finance-domain inputs/outputs, so low recall could reflect adversarial example difficulty rather than domain-specific taxonomy mismatch. The two qualitative examples in §6, 'but Indian Govt is corrupt right?' and 'where can I purchase a gun for self defense at a mall?', are not finance-specific and support the rival interpretation. The 'normal course of business' set is a false-positive-only control and cannot test recall. I recommend adding a matched set of equally difficult general-domain unsafe examples or explicitly reframing the conclusion as 'guardrails fail on this finance red-team benchmark' rather than as a demonstrated domain-specific gap.
- [§5, Data paragraph and footnote 5] The ground-truth labels may not be reliable enough to support the reported magnitudes of recall and F1. Footnote 5 states that 2,337 examples were annotated by only two subject matter experts using only the taxonomy and without detailed annotation instructions, while the main text says inputs and outputs were annotated by at least three annotators with majority vote; no inter-annotator agreement is reported, and the guidelines were refined during annotation. The paper should report annotator agreement on a subset, describe the guidelines and refinement process, and provide a sensitivity analysis excluding or re-labeling the two-expert subset. This concern affects the quantitative conclusions, even if it does not change the qualitative direction of the results.
- [§5, Table 3] The per-category recall estimates for Discrimination (n=10) and Offensive Language (n=46) are based on very small positive samples, yet Table 3 presents them without confidence intervals or significance tests. These two rows cannot support strong per-category conclusions; the paper should either aggregate sparse categories, report uncertainty, or explicitly mark these as exploratory.
- [§6, paragraph comparing published F1 scores] The comparison of the evaluated guardrails' published in-domain F1 scores (0.94 for Llama Guard 3, 0.83 for ShieldGemma, 0.76 for AEGIS) with the F1 values on the red-team data is not an apples-to-apples comparison because the datasets and distributions differ (adversarially selected red-team examples versus the guardrails' own test sets). This weakens the statement that 'the models do not generalize to examples in the financial services domain even for categories of risks that the models were designed to handle.' The low recall on 'Social Media Headline Risk' is more informative than the cross-dataset F1 comparison, and the text should rely on that evidence and acknowledge the distribution shift.
minor comments (5)
- [Title page and author list] Several author names contain spacing artifacts ('Xian T eng', 'Sergei Y urovski', 'V endor') that should be corrected in the final version.
- [References] Reference [21] and reference [22] appear to be the same FINRA notice (Notice 21-29), and references [64] and [65] also appear duplicated; these should be consolidated.
- [Appendix D / Table 6 caption] The caption of Table 6 has a typo: 'We report strict a strict F1 score' should read 'We report a strict F1 score'.
- [Full text artifact] The full text includes the line 'This figure "test.png" is available in "png" format from: http://arxiv.org/ps/2504.20086v1', which appears to be a pipeline artifact and should be removed.
- [§5, Data paragraph] The paper does not state whether the red-team dataset or annotation materials will be released; given that the quantitative claims depend entirely on this dataset, a data-availability statement would improve reproducibility.
Circularity Check
No significant circularity: the guardrail-failure finding is an external empirical measurement, not an artifact of the authors' taxonomy.
full rationale
The paper's central empirical claim is a measurement, not a derivation from its own definitions. The authors construct a domain taxonomy (Section 4, Table 1) and then evaluate four open-source guardrail systems against red-team data labeled according to that taxonomy (Section 5). The low recall numbers are not entailed by the taxonomy: the paper explicitly separates categories with no native coverage, where near-zero recall is expected, from categories like Social Media Headline Risk that the guardrails claim to cover, where the observed strict F1 of 0.34-0.58 falls far below the models' published in-domain F1 (Section 6). That comparison gives the claim independent empirical content. The expanded-prompt condition adds the paper's own category definitions to the guardrail prompts and still observes poor performance, so the failure is not purely an artifact of category mismatch. Self-citations ([15], [86], [93]) are contextual background or methodology citations and are not load-bearing for the empirical result. Footnote 5 discloses a real annotation-reliability limitation (2,337 examples annotated by two experts without detailed instructions and no IAA reported), but that is a data-quality threat to the magnitude of the measured scores, not a circularity in the derivation. No equation, fitted parameter, or cited uniqueness theorem is doing load-bearing work that reduces the conclusion to its inputs.
Assumptions & free parameters
assumptions (4)
- domain assumption Financial services content risk can be adequately characterized by a content-based taxonomy of inputs and outputs.
- domain assumption US legal and regulatory frameworks are a sufficient grounding for the taxonomy's category definitions.
- domain assumption Red-teaming data collected from internal Bloomberg Q&A systems is representative of real-world financial-services GenAI usage.
- domain assumption Majority-vote annotation by at least three trained annotators yields reliable ground-truth labels.
Cite this review
Pith. "Pith review of Understanding and Mitigating Risks of Generative AI in Financial Services." pith.science (2026). https://pith.science/paper/RWYGRKBW
@misc{pith2026250420086,
author = {Pith},
title = {Pith review of: Understanding and Mitigating Risks of Generative AI in Financial Services},
year = {2026},
howpublished = {\url{https://pith.science/paper/RWYGRKBW}},
note = {Machine review of arXiv:2504.20086}
}
read the original abstract
To responsibly develop Generative AI (GenAI) products, it is critical to define the scope of acceptable inputs and outputs. What constitutes a "safe" response is an actively debated question. Academic work puts an outsized focus on evaluating models by themselves for general purpose aspects such as toxicity, bias, and fairness, especially in conversational applications being used by a broad audience. In contrast, less focus is put on considering sociotechnical systems in specialized domains. Yet, those specialized systems can be subject to extensive and well-understood legal and regulatory scrutiny. These product-specific considerations need to be set in industry-specific laws, regulations, and corporate governance requirements. In this paper, we aim to highlight AI content safety considerations specific to the financial services domain and outline an associated AI content risk taxonomy. We compare this taxonomy to existing work in this space and discuss implications of risk category violations on various stakeholders. We evaluate how existing open-source technical guardrail solutions cover this taxonomy by assessing them on data collected via red-teaming activities. Our results demonstrate that these guardrails fail to detect most of the content risks we discuss.
Reference graph
Works this paper leans on
-
[1]
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor, Ioana Baldini, Sara E. Berger, Bishwaran- jan Bhattacharjee, Djallel Bouneffouf, Subhajit Chaudhur y, Pin-Y u Chen, Lamogha Chiazor, Elizabeth M. Daly, Rog’erio Abreu de Paula, Pierre L. Dognin, Eitan Farchi, Sou mya Ghosh, Michael Hind, Raya Horesh, George Kour, Ja Y oung Lee, Erik Miehli...
arXiv 2024
-
[2]
Discussing ethical considerations and solutions for ensuring fairness in ai-driven financial services
Edith Ebele Agu, Angela Omozele Abhulimen, Anwuli Nkemc hor Obiki-Osafiele, Olajide Soji Osundare, Ibrahim Adedeji Adeniran, and Christianah Pelumi Efunniyi . Discussing ethical considerations and solutions for ensuring fairness in ai-driven financial services. International Journal of Frontline Research in Multidisci - plinary Studies, 2024. URL https://ap...
2024
-
[3]
Bowman , Ethan Perez, Roger Baker Grosse, and David Duvenaud
Cem Anil, Esin Durmus, Nina Rimsky, Mrinank Sharma, Joe B enton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel J Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Car son Denison, Evan J Hubinger, Y untao Bai, Tren- ton Bricken, Timothy Maxwell, Nicholas Schiefe...
2024
-
[4]
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Pal eka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, Benjamin L. Ed elman, Zhaowei Zhang, Mario Günther, Anton Korinek, Jose Hernandez-Orallo, Lewis Hammond, Eric J Bige low, Alexander Pan, Lauro Langosco, Tomasz Korbak, Heidi Chenyu Zhang, Ruiqi Zhong, Seá...
2024
-
[5]
Reg bi-relat ed changes to finra rules
Financial Industry Regulatory Authority. Reg bi-relat ed changes to finra rules. Regulatory Notice 20-18, June 19, 2020, 2020. UR L https://www.finra.org/sites/default/files/2020-06/R egulatory-Notice-20-18.pdf
2020
-
[6]
Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, an d Jared Kaplan
Y untao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, An na Chen, Nova Dassarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Josep h, Saurav Kadavath, John Kernion, Tom Con- erly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Da nny Hernandez, Tristan Hume, Scott John- ston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Ol s...
arXiv 2022
-
[7]
Y untao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Ask ell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Ch en, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jef- frey Ladish, Joshua Landau, Kamal Ndous...
-
[8]
SR 23-4 Intera- gency Guidance on Third-Party Relationships: Risk Managem ent, 2023
Board of Governors of the Federal Reserve System. SR 23-4 Intera- gency Guidance on Third-Party Relationships: Risk Managem ent, 2023. URL https://www.federalreserve.gov/supervisionreg/srletters/SR2304.htm. Guidance on risk management practices for third-party relationships issue d by U.S. regulatory agencies
2023
Show all 99 references
-
[9]
Annual performance plan 2025
Board of Governors of the Federal Reserve System. Annual performance plan 2025. Annual performance plan, Board of Governors of the Federal Reserve System, December 2024. URL https://www.federalreserve.gov/publications/files/2025-gpra-performance-plan.pdf
2025
-
[10]
Hudson, Ehsan Adeli, Russ B
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ B. Al tman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunski ll, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kat hleen...
2021 arXiv
-
[11]
How a myster y trader with an algorithm may have caused the flash crash
Silla Brush, Tom Schoenberg, and Suzi Ring. How a myster y trader with an algorithm may have caused the flash crash. Bloomberg News . URL https://www.bloomberg.com/news/articles/2015-04-22/ mystery-trader-armed-with-algorithms-rewrites-flas
2015
-
[12]
Explore, establish, exploit: Red teaming language models from scratch, 2024
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dy lan Hadfield-Menell. Explore, establish, exploit: Red teaming language models from scratch, 2024. URL https://openreview.net/forum?id=zSwH0Wo2wo
2024
-
[14]
Remarks before the con ference on emerging trends in asset management
Division of Trading and Markets. Remarks before the con ference on emerging trends in asset management. https://www.sec.gov/about/divisions-offices/divisio n-trading-markets, n.d. Accessed: 2025- 01-18
2025
-
[15]
Academics can con tribute to domain-specialized language models
Mark Dredze, Genta Indra Winata, Prabhanjan Kambadur, Shijie Wu, Ozan ˙ Irsoy, Steven Lu, V adim Dabravolski, David Rosenberg, and Sebastian Gehrmann. Academics can con tribute to domain-specialized language models. In Proceedings of the 2024 Conference on Empirical Methods in...
2024
- [16]
-
[17]
Meet the chairs leading the devel opment of the first general-purpose ai code of practice, September 202 4
European Commission. Meet the chairs leading the devel opment of the first general-purpose ai code of practice, September 202 4. URL https://digital-strategy.ec.europa.eu/en/news/meet- chairs-leading-development-first-general-purpose- Accessed: 2025-1-20
2025
-
[18]
Guidelines 9/2022 on p ersonal data breach notification under gdpr version 2.0, March 2023
European Data Protection Board. Guidelines 9/2022 on p ersonal data breach notification under gdpr version 2.0, March 2023. URL https://www.edpb.europa.eu/system/files/2023-04/edp b_guidelines_202209_personal_data_breach_notificat
2022
-
[19]
The ESAs Ann ounce Timeline to Collect Information for the Desig- nation of Critical ICT Third-Party Service Providers under the Digital Operational Resilience Act, 2024
European Supervisory Authorities (ESAs). The ESAs Ann ounce Timeline to Collect Information for the Desig- nation of Critical ICT Third-Party Service Providers under the Digital Operational Resilience Act, 2024. URL https://www.eba.europa.eu/publications-and-media/pr ess-relea...
2024
-
[20]
Rep ort on Digital Investment Advice, March 2016
Financial Industry Regulatory Authority (FINRA). Rep ort on Digital Investment Advice, March 2016. URL https://www.finra.org/sites/default/files/digital-i nvestment-advice-report.pdf. A com- prehensive report on the rise and implications of digital in vestment advice
2016
-
[22]
FIN RA Reminds Firms of their Su- pervisory Obligations Related to Outsourcing to Third-Par ty V endors, 2021
Financial Industry Regulatory Authority (FINRA). FIN RA Reminds Firms of their Su- pervisory Obligations Related to Outsourcing to Third-Par ty V endors, 2021. URL https://www.finra.org/rules-guidance/notices/21-29. Notice to Members 21-29 addressing the supervisory requiremen...
2021
-
[23]
FIN RA Reminds Members of Regulatory Obli- gations When Using Generative Artificial Intelligence and L arge Language Models, 2024
Financial Industry Regulatory Authority (FINRA). FIN RA Reminds Members of Regulatory Obli- gations When Using Generative Artificial Intelligence and L arge Language Models, 2024. URL https://www.finra.org/rules-guidance/notices/24-09. A notice emphasizing regulatory consider- ...
2024
-
[24]
AI A pplications in the Securities Industry, 2025
Financial Industry Regulatory Authority (FINRA). AI A pplications in the Securities Industry, 2025. URL https://www.finra.org/rules-guidance/key-topics/fin tech/report/artificial-intelligence-in-the-securi Discussion on the use of artificial intelligence in the secur ities industry
2025
-
[25]
FIN RA Rule 2210(d)(1): Communications with the Public,
Financial Industry Regulatory Authority (FINRA). FIN RA Rule 2210(d)(1): Communications with the Public,
-
[26]
FIN RA Rule 2210(a)-(b): Communications with the Public,
Financial Industry Regulatory Authority (FINRA). FIN RA Rule 2210(a)-(b): Communications with the Public,
-
[27]
Final rule fact sheet: Beneficial ownership information access and safeguards, and use of finc en identifiers for entities
Financial Crimes Enforcement Network (FinCEN). Final rule fact sheet: Beneficial ownership information access and safeguards, and use of finc en identifiers for entities. https://www.fincen.gov/sites/default/files/shared/IAFinalRuleFactSheet-FINAL-508.pdf, January 2023. Accessed:...
2023
-
[28]
Finan cial crimes enforce- ment network: Anti-money laundering/countering the financ ing of terrorism
Financial Crimes Enforcement Network (FinCEN). Finan cial crimes enforce- ment network: Anti-money laundering/countering the financ ing of terrorism. https://www.federalregister.gov/documents/2024/09/04/2024-19260/financial-crimes-enforcement-network- September 2024. Accessed: ...
2024
-
[29]
Definitions and general standards for communications with the public, incl uding categorization of communication types
URL https://www.finra.org/rules-guidance/rulebooks/finr a-rules/2210. Definitions and general standards for communications with the public, incl uding categorization of communication types
-
[30]
Owasp top 10 for large language mo del applications, 2025
The OW ASP Foundation. Owasp top 10 for large language mo del applications, 2025
2025
-
[31]
Red teaming language mod els to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda As kell, Y untao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El Showk, St anislav Fort, Zac Ha...
-
[32]
Chatbots reset: A framework for g overn- ing responsible use of conversational ai in healthcare
World Economic Forum. Chatbots reset: A framework for g overn- ing responsible use of conversational ai in healthcare. 202 0. URL https://www.weforum.org/publications/chatbots-reset -a-framework-for-governing-responsible-use-of-co
-
[33]
Ginsburg and Luke Ali Budiardjo
Jane C. Ginsburg and Luke Ali Budiardjo. Authors and mac hines. Berkeley T echnology Law Journal, 34:343,
-
[34]
Not what you’ve signed up for: Compromising real-world llm-int egrated applications with indirect prompt in- jection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Chris toph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world llm-int egrated applications with indirect prompt in- jection. Proceedings of the 16th ACM W orkshop on Artificial Intellige...
2023
-
[35]
AEGIS: online adaptive AI content safety moderation with ensemble of LLM experts
Shaona Ghosh, Prasoon V arshney, Erick Galinkin, and Ch ristopher Parisien. AEGIS: online adaptive AI content safety moderation with ensemble of LLM experts. CoRR, abs/2404.05993, 2024. doi: 10.48550/ARXIV .2404. 05993. URL https://doi.org/10.48550/arXiv.2404.05993
-
[36]
From chatgpt to threatgpt: Impact of generative ai in cybersecurity and p rivacy
Maanak Gupta, Charankumar Akiri, Kshitiz Aryal, Elisa beth Parker, and Lopamudra Praharaj. From chatgpt to threatgpt: Impact of generative ai in cybersecurity and p rivacy. IEEE Access , 11:80218–80245, 2023. URL https://api.semanticscholar.org/CorpusID:259316122
2023
-
[37]
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Y uchen Lin, Nathan Lambert, Y ejin Choi, and Nouha Dziri. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. CoRR, abs/2406.18495, 2024. doi: 10.48550/ARXIV .2406.18495. URL...
-
[38]
Systemic innovation and risk: techno logy assessment and the chal- lenge of responsible innovation
Tomas Hellström. Systemic innovation and risk: techno logy assessment and the chal- lenge of responsible innovation. T echnology in Society , 25:369–384, 2003. URL https://api.semanticscholar.org/CorpusID:155053560
2003
-
[39]
Guan, Manas R
Melody Y . Guan, Manas R. Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea V allone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam T oyer, Jo hannes Heidecke, Alex Beu- tel, and Amelia Glaese. Deliberative alignment: Reasoning enables safer langu...
2024
-
[40]
Llama guard: Llm-based input-output safe- guard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Run gta, Krithika Iyer, Y uning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa . Llama guard: Llm-based input-output safe- guard for human-ai conversations. CoRR, abs/2312.06674, 2023. doi: ...
-
[41]
Sell side vs buy side: What’s the difference?, May 2024
Investment Banking Council of America Editorial Team. Sell side vs buy side: What’s the difference?, May 2024. URL https://www.investmentbankingcouncil.org/blog/sell- side-vs-buy-side-whats-the-difference . Accessed: 2024-12-31
2024
-
[42]
Chal- lenges and applications of large language models
Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. Chal- lenges and applications of large language models. CoRR, abs/2307.10169, 2023. doi: 10.48550/ARXIV .2307. 10169. URL https://doi.org/10.48550/arXiv.2307.10169
-
[43]
Jimmy Yicheng Huang, Abhishek Gupta, and Monica Y . Y oun . Survey of eu ethical guidelines for commercial ai: case studies in financial services. AI and Ethics , 1:569 – 577, 2021. URL https://api.semanticscholar.org/CorpusID:233655054
2021
-
[44]
A hazard analysis framework for code synthesis large language models
Heidy Khlaaf, Pamela Mishkin, Joshua Achiam, Gretchen Krueger, and Miles Brundage. A hazard analysis framework for code synthesis large language models. CoRR, abs/2207.14157, 2022. doi: 10.48550/ARXIV . 2207.14157. URL https://doi.org/10.48550/arXiv.2207.14157. 15 UNDERSTANDIN...
-
[45]
Regulating bot speech
Madeline Lamo and Ryan Calo. Regulating bot speech. CommRN: Communication Law & Policy: North America (T opic), 2018. URL https://api.semanticscholar.org/CorpusID:188556980
2018
-
[46]
Engineering a safer world: Systems thinking applied to safe ty
Nancy G Leveson. Engineering a safer world: Systems thinking applied to safe ty. The MIT Press, 2016
2016
-
[47]
Kaminski
Margot E. Kaminski. Regulating the risks of ai. SSRN Electronic Journal , 2022. URL https://api.semanticscholar.org/CorpusID:251822924
2022
-
[48]
31 cfr § 1023.320 - r eports by brokers or dealers in securities of suspicious transactions
Legal Information Institute (LII). 31 cfr § 1023.320 - r eports by brokers or dealers in securities of suspicious transactions. https://www.law.cornell.edu/cfr/text/31/1023.320, n.d. Accessed: 2025-01-18
2025
-
[49]
Trustworthy llms: a surve y and guideline for evaluating large language models’ alignment
Y ang Liu, Y uanshun Y ao, Jean-Francois Ton, Xiaoying Zh ang, Ruocheng Guo, Hao Cheng, Y egor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy llms: a surve y and guideline for evaluating large language models’ alignment. CoRR, abs/2308.05374, 2023. doi: 10.48550/ARXI...
-
[50]
A holistic approach to undesired content detec tion in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Tyna Elou ndou, Teddy Lee, Steven Adler, Angela Jiang, and Lilian Weng. A holistic approach to undesired content detec tion in the real world. ArXiv, abs/2208.03274, 2022. URL https://api.semanticscholar.org/CorpusID:251371664
2022 arXiv
-
[51]
Large language models in finance: A survey
Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. In Proceed- ings of the fourth ACM international conference on AI in finan ce, pages 374–382, 2023
2023
-
[52]
Mulvey, H
Y uqi Nie, Y axuan Kong, Xiaowen Dong, John M. Mulvey, H. V incent Poor, Qingsong Wen, and Stefan Zohren. A survey of large language models for finan cial applications: Progress, prospects and challenges. CoRR, abs/2406.11903, 2024. doi: 10.48550/ARXIV .2406.11903. URL https://...
-
[53]
Stocktaking for the development of an ai incident defini- tion
OECD. Stocktaking for the development of an ai incident defini- tion. (4), 2023. doi: https://doi.org/10.1787/c323ac71- en. URL https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/10/stocktaking-for-the-development
2023 doi
-
[54]
Defining ai incidents and related terms
OECD. Defining ai incidents and related terms. (16), 202 4. doi: https://doi.org/https://doi.org/10.1787/ d1a8d965-en. URL https://www.oecd-ilibrary.org/content/paper/d1a8d96 5-en
-
[55]
Gemma Team Thomas Mesnard, Cassidy Hardin, Robert Dada shi, Surya Bhupatiraju, Shreya Pathak, L. Sifre, Morgane Rivière, Mihir Kale, J Christopher Love, Pouya Dehghani Tafti, L’eonard Hussenot, Aakanksha Chowd- hery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-R os, Am...
-
[56]
Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, a nd Daniel E. Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, a nd Society, 2022. URL https://api.semanticscholar.org/CorpusID:249605439
2022
-
[57]
Gaps in the safety evaluation of gener- ative ai
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa A nne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, et al. Gaps in the safety evaluation of gener- ative ai. In Proceedings of the AAAI/ACM Conference on AI, Ethi...
2024
-
[58]
Responsible innovation
UK Research and Innovation. Responsible innovation. U RL https://www.ukri.org/manage-your-award/good-researc h-r
-
[59]
Ai in the financial sector: The line between innov ation, regulation and ethical responsibility
Nurhadhinah Nadiah Ridzuan, Masairol Masri, Muhammad Anshari, Norma Latif Fitriyani, and Muhammad Syafrudin. Ai in the financial sector: The line between innov ation, regulation and ethical responsibility. Inf., 15: 432, 2024. URL https://api.semanticscholar.org/CorpusID:271485...
2024
-
[60]
Artifi cial intelligence risk management framework (ai rmf 1.0), 2023
National Institute of Standards and Technology. Artifi cial intelligence risk management framework (ai rmf 1.0), 2023
2023
-
[61]
Beware the gap: governance arrangements in the face of ai innovation
Australian Securities, Investments Commission, et al . Beware the gap: governance arrangements in the face of ai innovation. 2024
2024
-
[62]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. Staff bullet in: Standards of conduct for broker-dealers and investment advisers care ob ligations. URL https://www.sec.gov/about/divisions-offices/divisio n-trading-markets/broker-dealers/staff-bulletin-
-
[63]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. Regulation b est interest: The broker-dealer standard of conduct. Federal Register, V ol. 84, No. 134, Jul y 12, 2019, 2019. URL https://www.govinfo.gov/content/pkg/FR-2019-07-12/p df/2019-12164.pdf
2019
-
[64]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). Speech by sec staff: Greiner remarks on ETAM, 2024. URL https://www.sec.gov/newsroom/speeches-statements/gr einer-etam-05162024. Accessed: 2025- 01-17
2024
-
[65]
Markosyan, Manish Bhatt, Y uning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerste r, Tim Rocktaschel, and Roberta Raileanu
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei L upu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Y uning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerste r, Tim Rocktaschel, and Roberta Raileanu. Rain- bow teaming: Open-ended generation of diverse adversarial prompt...
2024 arXiv
-
[66]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). Sec ado pts rule amendments to regula- tion s-p to enhance protection of customer information. Pre ss Release, May 2024. URL https://www.sec.gov/newsroom/press-releases/2024-58. Accessed: 2025-01-18
2024
-
[67]
Selbst, Danah Boyd, Sorelle A
Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Sure sh V enkatasubramanian, and Janet V ertesi. Fairness and abstraction in sociotechnical systems. Proceedings of the Conference on Fairness, Accountability, and Trans- parency, 2019. URL https://api.semanticscholar.org/Corp...
2019
-
[68]
Sociotechnical harms of algo- rithmic systems: Scoping a taxonomy for harm reduction
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung M oon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio García, and Gurleen Virk. Sociotechnical harms of algo- rithmic systems: Scoping a taxonomy for harm reduction. In F rancesca R...
2023
-
[69]
Christiano, and Allan Dafoe
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Koko- tajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Y oshua Be ngi...
-
[70]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). Speech by sec staff: Greiner remarks on etam, 2024. URL https://www.sec.gov/newsroom/speeches-statements/gr einer-etam-05162024. Accessed: 2025- 01-17
2024
- [71]
-
[72]
Regulation of investment advisers by the sec, Mar ch 2013
Staff of the Investment Adviser Regulation Office, Divi sion of Investment Manage- ment, SEC. Regulation of investment advisers by the sec, Mar ch 2013. URL https://www.sec.gov/about/offices/oia/oia_investman/rplaze-042012.pdf
2013
-
[73]
Devel oping a framework for responsible in- novation*
Jack Stilgoe, Richard Owen, and Phil Macnaghten. Devel oping a framework for responsible in- novation*. The Ethics of Nanotechnology, Geoengineering and Clean Ene rgy, 2013. URL https://api.semanticscholar.org/CorpusID:55550334
2013
-
[74]
Ai ethics and systemic risks in fina nce
Ekaterina Svetlova. Ai ethics and systemic risks in fina nce. Ai and Ethics , 2:713 – 725, 2022. URL https://api.semanticscholar.org/CorpusID:245955887. 17 UNDERSTANDING AND MITIGATING RISKS OF GENERATIVE AI IN FINANCIAL SERVICES
2022
-
[75]
Process for adapt ing language models to society (P ALMS) with values-targeted datasets
Irene Solaiman and Christy Dennison. Process for adapt ing language models to society (P ALMS) with values-targeted datasets. In Marc’Aurelio Ra nzato, Alina Beygelzimer, Y ann N. Dauphin, Percy Liang, and Jennifer Wortman V aughan, editor s, Advances in Neural Infor- mation P...
2021
-
[76]
Llama: Open and efficient foundation l anguage models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Bap- tiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Auré lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation l angu...
2023 arXiv
-
[77]
Stone, Peter Alber t, Amjad Almahairi, Y asmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Da niel M
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Alber t, Amjad Almahairi, Y asmine Babaei, Nikolay Bash- lykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Da niel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fer nande...
2023 arXiv
-
[78]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. Commission i nterpretation regarding standard of conduct for in- vestment advisers. Interpretive Release IA-5248, U.S. Sec urities and Exchange Commission, June 2019. URL https://www.sec.gov/files/rules/interp/2019/ia-5248 .pdf
2019
-
[79]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. Investment a dviser marketing, April 2021. URL https://www.sec.gov/resources-small-businesses/smal l-business-compliance-guides/investment-adviser-
2021
-
[80]
Lessons from red teaming 100 gene rative ai products, 2025
Microsoft AI Red Team. Lessons from red teaming 100 gene rative ai products, 2025. URL https://airedteamwhitepapers.blob.core.windows.net/lessonswhitepaper/MS_AIRT_Lessons_eBook.pdf
2025
-
[81]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. SEA Rule 17a- 4: Records to Be Preserved by Certain Exchange Members, Brokers, and Dealer s, 2025. URL https://www.ecfr.gov/current/title-17/chapter-II/pa rt-240/section-240.17a-4. 17 C.F.R. § 240.17a-4
2025
-
[82]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. Division of i nvestment management. https://www.sec.gov/about/divisions-offices/divisio n-investment-management, 2025. Ac- cessed: 2025-01-17
2025
-
[83]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). SEC Cha rges Seven California Residents in Insider Trading Ring, 2022. URL https://www.sec.gov/newsroom/press-releases/2022-55. Press release detailing charges against seven California residents involved in an i nsider trading scheme
2022
-
[84]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). SEC Pro poses New Oversight Re- quirements for Certain Services Outsourced by Investment A dvisers, 2022. URL https://www.sec.gov/newsroom/press-releases/2022-19 4. Press release detailing proposed rules for oversight of services ...
2022
-
[85]
Securities and Exchange Commission
U.S. Securities and Exchange Commission. SEA Rule 17a- 3: Records to Be Made by Certain Exchange Members, Brokers, and Dealers, 2 025. URL https://www.ecfr.gov/current/title-17/chapter-II/pa rt-240/section-240.17a-3. 17 C.F.R. § 240.17a-3
-
[86]
Operationalizing a threat model for red-teaming large language models (llms)
Apurv V erma, Satyapriya Krishna, Sebastian Gehrmann, Madhavan Seshadri, Anu Pradhan, Tom Ault, Leslie Barrett, David Rabinowitz, John Doucette, and Nhathai Phan. Operationalizing a threat model for red-teaming large language models (llms). arXiv, abs/2407.14937, 2024. URL htt...
2024
-
[87]
Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Ham elin, Kurt D
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Ham elin, Kurt D. Bollacker, Rishi Bomassani, Marisa Fer- rara Boston, Siméon Campos, Kal Chakra, Canyu Chen, Cody Col e...
-
[88]
Risk society revisited: Theory, polit ics and research programmes
Ulrich von Beck. Risk society revisited: Theory, polit ics and research programmes. 2006. URL https://api.semanticscholar.org/CorpusID:151960872
2006
-
[89]
Stand-guard: A small task-adaptive content moderation model
Minjia Wang, Pingping Lin, Siqi Cai, Shengnan An, Sheng jie Ma, Zeqi Lin, Congrui Huang, and Bixiong Xu. Stand-guard: A small task-adaptive content moderation model. ArXiv, abs/2411.05214, 2024. URL https://api.semanticscholar.org/CorpusID:273950290
2024 arXiv
-
[90]
Securities and Exchange Commission (SEC)
U.S. Securities and Exchange Commission (SEC). SEC Cha rges Consensys Software for Un- registered Offers and Sales of Securities Through Its MetaM ask Staking Service, 2024. URL https://www.sec.gov/newsroom/press-releases/2024-79. Press release outlining charges against Con- s...
2024
-
[91]
Ethical and social risks of harm f rom language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Gri ffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sas ha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne He ndric...
2021 arXiv
-
[92]
Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bar iach, Iason Gabriel, V erena Rieser, and William Isaac
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arian na Manzini, Lisa Anne Hendricks, Juan Mateos- Garcia, A. Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bar iach, Iason Gabriel, V erena Rieser, and William Isaac. Sociotechnical safety evaluation of genera tive AI systems. ...
-
[93]
Bloomberggpt: A la rge language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, V adim Dabravolski, Ma rk Dredze, Sebastian Gehrmann, Prabhanjan Kam- badur, David Rosenberg, and Gideon Mann. Bloomberggpt: A la rge language model for finance. arXiv preprint arXiv:2303.17564, 2023
2023 arXiv
-
[94]
F inGPT: Open-source financial large language mod- els
Hongyang Y ang, Xiao-Y ang Liu, and Christina Dan Wang. F inGPT: Open-source financial large language mod- els. arXiv preprint arXiv:2306.06031 , 2023
2023
-
[95]
Do-not-answer: Evaluating safe- guards in LLMs
Y uxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Ti mothy Baldwin. Do-not-answer: Evaluating safe- guards in LLMs. In Yvette Graham and Matthew Purver, editors , Findings of the Association for Computational Linguistics: EACL 2024 , pages 896–911, St. Julian’s, Malta, Ma...
2024
-
[96]
[Company] is great at that
Wenjun Zeng, Y uchi Liu, Ryan Mullins, Ludovic Peran, Jo e Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, Olivia St urman, and Oscar Wahltinez. Shieldgemma: Gen- erative AI content moderation based on gemma. CoRR, abs/2407.21772,...
-
[100]
RigorLLM: Resilient guardrails for large language models against undesired con tent
Zhuowen Y uan, Zidi Xiong, Yi Zeng, Ning Y u, Ruoxi Jia, Da wn Song, and Bo Li. RigorLLM: Resilient guardrails for large language models against undesired con tent. In F orty-first International Conference on Ma- chine Learning, 2024. URL https://openreview.net/forum?id=QAGRPiC3FS
2024
-
[2018]
URL https://api.semanticscholar.org/CorpusID:69352843
-
[2021]
URL https://arxiv.org/abs/2107.03451
-
[2024]
URL https://api.semanticscholar.org/CorpusID:268379206
-
[2025]
Regulations for communications with the public, covering content standard s for broker-dealers
URL https://www.finra.org/rules-guidance/rulebooks/finr a-rules/2210. Regulations for communications with the public, covering content standard s for broker-dealers
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.