A seven-level ordinal severity scale for tool-using AI agents, computed from execution traces, reveals cases where binary attack-success-rate metrics hide dangerous cross-scope leaks and worsening tail risk.
Judging LLM-as-a-judge with MT-Bench and chatbot arena,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
A seven-level ordinal severity scale for tool-using AI agents, computed from execution traces, reveals cases where binary attack-success-rate metrics hide dangerous cross-scope leaks and worsening tail risk.