REVIEW 2 cited by
SecretBench: A Dataset of Software Secrets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
According to GitGuardian's monitoring of public GitHub repositories, the exposure of secrets (API keys and other credentials) increased two-fold in 2021 compared to 2020, totaling more than six million secrets. However, no benchmark dataset is publicly available for researchers and tool developers to evaluate secret detection tools that produce many false positive warnings. The goal of our paper is to aid researchers and tool developers in evaluating and improving secret detection tools by curating a benchmark dataset of secrets through a systematic collection of secrets from open-source repositories. We present a labeled dataset of source codes containing 97,479 secrets (of which 15,084 are true secrets) of various secret types extracted from 818 public GitHub repositories. The dataset covers 49 programming languages and 311 file types.
Forward citations
Cited by 2 Pith papers
-
Extended Version: It Should Be Easy but... New Users Experiences and Challenges with Secret Management Tools
New users of secret management tools can store and read secrets fairly easily, but injecting them into an app is much harder because official documentation omits relevant examples and flag explanations, driving users ...
-
Secret Breach Detection in Source Code with Large Language Models
Fine-tuned open-source LLMs classify regex-extracted secret candidates with 0.985 binary F1 and 0.982 multiclass F1 on the SecretBench benchmark.
Discussion (0). Continue with ORCID to comment.