Pith. sign in

REVIEW 4 cited by

On the (In)Security of LLM App Stores

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.08422 v2 pith:YTSOP5UX submitted 2024-07-11 cs.CR cs.AI

classification cs.CRcs.AI
keywords appspotentialsecuritystorescollectedidentifymaliciousabusive
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

LLM app stores have seen rapid growth, leading to the proliferation of numerous custom LLM apps. However, this expansion raises security concerns. In this study, we propose a three-layer concern framework to identify the potential security risks of LLM apps, i.e., LLM apps with abusive potential, LLM apps with malicious intent, and LLM apps with exploitable vulnerabilities. Over five months, we collected 786,036 LLM apps from six major app stores: GPT Store, FlowGPT, Poe, Coze, Cici, and Character.AI. Our research integrates static and dynamic analysis, the development of a large-scale toxic word dictionary (i.e., ToxicDict) comprising over 31,783 entries, and automated monitoring tools to identify and mitigate threats. We uncovered that 15,146 apps had misleading descriptions, 1,366 collected sensitive personal information against their privacy policies, and 15,996 generated harmful content such as hate speech, self-harm, extremism, etc. Additionally, we evaluated the potential for LLM apps to facilitate malicious activities, finding that 616 apps could be used for malware generation, phishing, etc. Our findings highlight the urgent need for robust regulatory frameworks and enhanced enforcement mechanisms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Large-Scale Empirical Analysis of Custom GPTs' Vulnerabilities in the OpenAI Ecosystem

    cs.CR 2025-05 conditional novelty 6.0 of 10

    More than 95% of tested custom GPTs in the OpenAI store are vulnerable to at least one of seven jailbreak attacks, and the paper proposes a multi-metric ranking to link popularity with security risk.

  2. Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers

    cs.HC 2024-11 conditional novelty 6.0 of 10

    Most professional artists lack the awareness and ability to use robots.txt, many hosting platforms do not let them edit it, and a substantial share of AI assistant crawlers ignore the protocol.

  3. Tracking GPTs Third Party Service: Automation, Analysis, and Insights

    cs.CR 2025-06 conditional novelty 4.0 of 10

    GPTs-ThirdSpy uses GUI automation to extract third-party service domains and privacy policy links from GPT Store settings pages, yielding 109 domains across 500 popular GPTs.

  4. RedTeamLLM: an Agentic AI framework for offensive security

    cs.CR 2025-05 conditional novelty 3.0 of 10

    The paper reports that adding a separate reasoning step to a terminal-operating LLM agent reduces tool calls and improves completion on 4 of 5 entry-level CTF virtual machines, while the framework's memory and plan-co...

Pith tools