Pith. sign in

REVIEW 1 cited by

LLMs Perform Poorly at Concept Extraction in Cyber-security Research Literature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07110 v1 pith:SFUGS7ZD submitted 2023-12-12 cs.CL cs.CRcs.LG

LLMs Perform Poorly at Concept Extraction in Cyber-security Research Literature

classification cs.CL cs.CRcs.LG
keywords domainllmscybersecurityresultssometrendsentitiesextract
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The cybersecurity landscape evolves rapidly and poses threats to organizations. To enhance resilience, one needs to track the latest developments and trends in the domain. It has been demonstrated that standard bibliometrics approaches show their limits in such a fast-evolving domain. For this purpose, we use large language models (LLMs) to extract relevant knowledge entities from cybersecurity-related texts. We use a subset of arXiv preprints on cybersecurity as our data and compare different LLMs in terms of entity recognition (ER) and relevance. The results suggest that LLMs do not produce good knowledge entities that reflect the cybersecurity context, but our results show some potential for noun extractors. For this reason, we developed a noun extractor boosted with some statistical analysis to extract specific and relevant compound nouns from the domain. Later, we tested our model to identify trends in the LLM domain. We observe some limitations, but it offers promising results to monitor the evolution of emergent trends.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs

    cs.CL 2025-10 conditional novelty 5.0

    An LLM-based pipeline extracts technology triples from arXiv and patent full text, tracks rising topic co-occurrence, and labels retrieval-augmented generation and conversational agents as emerging transformative tech...