pith. sign in

Theoremqa: A theorem-driven question answering dataset

10 Pith papers cite this work. Polarity classification is still indexing.

10 Pith papers citing it

citation-role summary

dataset 2 background 1 other 1

citation-polarity summary

years

2026 8 2023 2

representative citing papers

Detecting Language Model Attacks with Perplexity

cs.CL · 2023-08-27 · unverdicted · novelty 5.0

Jailbreak prompts with adversarial suffixes have high GPT-2 perplexity, and a LightGBM model on perplexity and length detects most attacks.

citing papers explorer

Showing 10 of 10 citing papers.