Pith. sign in

REVIEW 11 cited by

Granite Guardian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07724 v2 pith:QDCTYHPL submitted 2024-12-10 cs.CL

classification cs.CL
keywords graniteguardianmodelsriskacrosscontentdetectionmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with any large language model (LLM). These models offer comprehensive coverage across multiple risk dimensions, including social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and hallucination-related risks such as context relevance, groundedness, and answer relevance for retrieval-augmented generation (RAG). Trained on a unique dataset combining human annotations from diverse sources and synthetic data, Granite Guardian models address risks typically overlooked by traditional risk detection models, such as jailbreaks and RAG-specific issues. With AUC scores of 0.871 and 0.854 on harmful content and RAG-hallucination-related benchmarks respectively, Granite Guardian is the most generalizable and competitive model available in the space. Released as open-source, Granite Guardian aims to promote responsible AI development across the community. https://github.com/ibm-granite/granite-guardian

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers

    cs.AI 2026-05 conditional novelty 7.0 of 10

    Typed quantity verification exposes a canonical-equivalence blind spot in neural fact-checkers; Symbolic Augmentation fixes it (36.5%→98.2%) and transfers to SciFact-Open (+0.037 binary macro-F1).

  2. JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A guard trained to anticipate safety-relevant futures from partial trajectories cuts average attack success from 23.0% to 7.1% across four agent-safety benchmarks.

  3. DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

    cs.CR 2026-07 conditional novelty 6.0 of 10

    An evolving attack-defense loop, DARWIN, achieves state-of-the-art jailbreak success rates on frontier LLMs/guardrails and trains a guardrail with 91.6% average unsafe recall while retaining ~100% benign pass rate.

  4. CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization

    cs.CR 2026-07 conditional novelty 6.0 of 10

    CPInj demonstrates that federated textual prompt optimization (a TextGrad-style loop) is vulnerable to a multi-objective injection attack that persists through aggregation, degrades accuracy by up to 55 points, and ou...

  5. HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A hypernetwork maps layer-wise activation fingerprints of a fine-tuned LLM to a Safe Side Network that routes harmful prompts to refusal without editing model weights.

  6. Adversarial Bug Reports as a Security Risk in Language Model-Based Automated Program Repair

    cs.SE 2025-09 conditional novelty 6.0 of 10

    Adversarial bug reports induced attacker-desired patches in 90% of trials, while the best tested pre-repair filter caught only 47%, exposing a structural weakness in LLM-based automated program repair.

  7. Concealment of Intent: A Game-Theoretic Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Intent-hiding adversarial prompting that mixes malicious intents with innocuous skills bypasses prompt and response filters, and a game-theoretic analysis quantifies the attacker's scaling advantage.

  8. Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation

    cs.CL 2026-04 conditional novelty 5.5 of 10

    Qwen Guard (4B) reaches 83.97% recall on a 79k NIST-aligned safety benchmark while larger models such as Llama Guard 12B and GPT-OSS 20B miss up to 75% of unsafe content; model size does not predict detection performance.

  9. Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

    cs.LG 2026-05 reject novelty 5.0 of 10

    A 184M-parameter DeBERTa-v3 fine-tuned model is claimed to beat Llama-Guard-3-8B on all tested prompt-injection benchmarks while adding BFSI regulatory labels, but a leaked training/eval overlap undermines the zero-FPR claim.

  10. SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues

    cs.CL 2025-05 conditional novelty 5.0 of 10

    STREAM fine-tunes a small reasoning model on human-labeled, reason-annotated multi-turn dialogues and uses it to warn target LLMs, cutting average attack success rates by roughly half while keeping benchmark scores close.

  11. JavelinGuard: Low-Cost Transformer Architectures for LLM Security

    cs.LG 2025-06 reject novelty 4.0 of 10

    A study of five small transformer classifier architectures for LLM jailbreak and prompt injection detection claims low-latency accuracy comparable to large models, led by the multi-task Raudra design.

Pith tools