Pith. sign in

REVIEW 4 major objections 6 minor

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

T0 review · 4 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Commercial AI system prompts often protect users only shallowly, and about 40% still contain instructions that work against them.

desk verdict Solid first comparative audit of commercial system prompts: usable taxonomy, real org/time patterns, main caveat is leaked-corpus provenance already flagged by the authors. read the letter →

arxiv 2607.28617 v2 pith:YDN6QNZ3 submitted 2026-07-30 cs.AI cs.CLcs.CYcs.HC

classification cs.AIcs.CLcs.CYcs.HC
keywords systempromptsAIauditinguserprotectionLLMapplicationstransparencypromptgovernanceAISPAcommercial
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Hidden system prompts steer how commercial AI products behave before any user speaks, yet they are rarely disclosed or independently checked. This paper introduces AISPA, an eight-dimension, user-centered audit that labels individual instructions as protective or problematic. Applied to 3,249 instructions across 88 real products, the audit finds protection is nearly universal but thin: almost every product has at least one protective line, yet only about a quarter cover all eight dimensions. Prompts have grown longer and more protective over time, but roughly two in five products still include at least one instruction that undermines user interests, and protective and harmful directives often sit in the same prompt. Organizations differ sharply in how carefully they write these rules. The authors argue this layer needs transparency standards and third-party oversight without forcing full public disclosure.

What carries the argument

AISPA: a span-level audit taxonomy of eight user-rights dimensions (identity transparency, truthfulness, privacy, tool/action safety, user agency, unsafe-request handling, harm prevention, fairness/inclusion/neutrality), scoring each auditable instruction +1 protective or −1 problematic via a three-round human–LLM pipeline.

What would settle it

Obtain official current system prompts from a large, stratified sample of the same products and re-run the AISPA labeling; if comprehensive eight-dimension coverage rises well above ~24% and products with any problematic instruction fall well below ~40%, the prevalence claims fail.

Watch

Extended reading notes

Core claim

Across 88 commercial AI products, protective system-prompt instructions are near-universal (98.9% of products have at least one) but shallow (only about 24% cover all eight AISPA dimensions), while roughly 40% of products contain at least one problematic instruction that works against user interests, with large organization-level gaps and frequent coexistence of protective and problematic text in the same prompt.

Load-bearing premise

The leaked or community-disclosed prompts from public repositories are authentic and representative enough of live commercial products to support product- and organization-level prevalence claims.

Editorial extensions

If this is right

  • Third-party pre-deployment prompt certification becomes a concrete accountability tool without requiring full public disclosure.
  • Product and organization rankings on protective versus problematic counts can become a public quality signal for users and regulators.
  • Prompt design standards can target the thinnest dimensions (notably privacy and unsafe-request handling) and the gray-area patterns of identity concealment and parasocial dependency.
  • Temporal growth in protective instructions can be tracked as an industry norm rather than left to ad hoc developer practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If certification status is public while prompt text stays private, market pressure may close the org-level gap faster than regulation alone.
  • Gray-area instructions (human mimicry, override keys, content-policy relaxations) may become the next contested frontier once binary problematic cases decline.
  • Specialized and agentic products may need dimension weights different from general chatbots, because autonomy–agency tradeoffs are structural there.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces AISPA, an eight-dimension, UDHR-anchored taxonomy and a three-round human–LLM workflow for span-level auditing of commercial LLM system prompts. Auditors label non-core-logic spans as protective (+1) or problematic (−1) along identity transparency, truthfulness, privacy, tool/action safety, user agency, unsafe-request handling, harm prevention, and fairness. Applying this protocol to leaked/disclosed prompts from 88 products, the authors report that protective instructions are near-universal but shallow (98.9% of products have ≥1; only ~24% cover all eight dimensions), that prompts have lengthened and grown more protective over 2024–2025, that organization-level protective density varies widely (e.g., Anthropic vs. weaker peers), and that roughly two-fifths of products still contain at least one user-adverse instruction, often coexisting with protective text. A separate gray-area analysis describes borderline design patterns (identity mimicry, parasocial cues, permission overrides, unrestricted content policies).

Significance. System prompts are a consequential and under-scrutinized control layer in deployed LLM products; a reusable user-centric audit codebook plus the first multi-product empirical map is a genuine contribution to AI governance and HCI/safety practice. Strengths include a clear unit of analysis (spans), explicit core vs. non-core scope rules, high calibration IAA (0.933), an asymmetric unanimous-expert bar for −1 labels, temporal and provider case studies, an honest limitations section, and Appendix B authenticity checks (maintainer contact and cross-repository Dice overlap). If the descriptive prevalence patterns hold under clearer sampling caveats, the work supplies concrete evidence for transparency, standardization, and third-party prompt review—without requiring full public disclosure of proprietary prompts.

major comments (4)
  1. [Abstract; §5.1] Abstract and opening claim “3,249 instructions,” but §5.1 reports a final dataset of 2,420 entries (2,346 protective + 74 problematic) from 1,818 spans, plus 44 gray-area entries. This is a load-bearing numerical inconsistency for the paper’s headline audit scale. Please reconcile the abstract, §1/§8, and §5.1 (e.g., candidate spans vs. retained entries vs. gray-area) and use one consistent accounting everywhere.
  2. [§5.2.2; Figure 1; Figure 7] Finding 1 and Figures 1/7 rank organizations by average protective/problematic counts, but many organizations appear to contribute a single product (or very few). With n_org often ≈1, “organization averages” are product point estimates and are sensitive to category mix (chatbot vs. coding agent). Report n per organization, avoid over-interpreting single-product ranks as institutional policy, and consider category-stratified or mixed-effects summaries before strong org-level claims.
  3. [Abstract; §5.1; Limitations; Appendix B] The central prevalence claims (98.9% ≥1 protective; ~38.6% ≥1 problematic; shallow eight-dimension coverage) are generalized to “commercial AI products,” yet the corpus is leaked/community-disclosed prompts with acknowledged selection toward extractable systems (Limitations; Appendix B: 38/88 lack cross-repo matches). The validation work is real but incomplete for product- and market-level inference. Tighten abstract/conclusion wording to “in this leaked corpus,” quantify how sensitive headline percentages are to excluding low-overlap or single-source prompts, and state what claims remain if the sample is biased toward longer or more jailbreak-exposed prompts.
  4. [§5.2.1; Figure 5] Figure 5’s temporal story (longer, more protective prompts; fluctuating problematic rates) bins heterogeneous products with small per-bin n (e.g., 2025-Q1 n=9) and does not control for shifting category composition or uncertain leak dates vs. deployment dates. Composition change could mimic “industry learning.” Either restrict the panel to repeated product lines (as in the Claude/GPT/Grok case study in Figure 8) or show category-adjusted trends and uncertainty; otherwise soften causal language about protection “becoming a more visible concern.”
minor comments (6)
  1. [Abstract; §5.2.1] Abstract says “only 24%” and “roughly 40%” while §5.2.1 gives 23.9% and 38.6%; keep one precise figure and use “approximately” consistently.
  2. [Figure 6; §5.2.1] Figure 6(b) percentages for problematic D5/D2 (18% / 15% in the plot text vs. 18.2% / 14.8% in prose) should be aligned and include absolute counts.
  3. [§4.1; §5.1] Define “entry” vs. “span” vs. “instruction” once early (footnote 8 helps) and use it uniformly in figures and the abstract.
  4. [Table 1; §9] Table 1 examples are effective; briefly note whether quotes are lightly redacted and whether products are identifiable from the full audit release plan.
  5. [§7] Related Work (§7) is thin on prompt-injection/hardening and industry model-spec literature; a short contrast paragraph would better situate AISPA’s user-protective (not system-defensive) stance.
  6. [§8; title page] Minor prose issues: “comprises of” → “comprises” (§8); ensure arXiv/author URL formatting is consistent in the header.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: observational span-level audit with predefined taxonomy applied to external prompts.

full rationale

AISPA is an empirical content-analysis paper, not a first-principles derivation. The load-bearing claims are descriptive prevalence statistics (e.g., 98.9% of products have ≥1 protective entry; 23.9% cover all eight dimensions; ~38.6% have ≥1 problematic entry; org-level averages such as Anthropic 62.3 protective / 0.1 problematic) obtained by applying a pre-specified eight-dimension codebook and +1/−1 polarity rules to leaked system-prompt text from 88 products. The taxonomy is defined before the audit (Section 3; Table 1; UDHR anchors), spans are external artifacts, labels come from a three-round human–LLM protocol with reported IAA (0.933) and unanimous-expert threshold for −1, and time/org trends are counts over that labeled corpus—not fitted parameters re-exported as predictions. Related-work self-citations (e.g., co-author safety/agent papers) are background and do not force the prevalence results. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz-via-citation chain appears in the claimed findings. Theory-ladenness of ‘protective’ vs ‘problematic’ is a construct-validity issue, not circularity by construction. Score 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 3 invented entities

Load-bearing structure is normative and methodological rather than parametric. Claims rest on (i) treating system-prompt text as a primary behavioral governor, (ii) UDHR-mapped user-interest dimensions as the right audit axes, (iii) span-level ±1 coding as a valid proxy for user protection, and (iv) leaked corpus authenticity/representativeness. No physical constants or curve-fit parameters drive the headlines; free choices are taxonomy design and adjudication thresholds.

free parameters (3)
  • Unanimous three-expert threshold for retaining −1 labels = 3/3 experts must agree on −1
    Asymmetric adjudication rule that directly shapes the reported problematic prevalence; chosen for reputational caution, not estimated from external outcome data.
  • Eight-dimension AISPA partition and span merge rules = 8 dimensions; sentence-default spans
    Hand-designed taxonomy and ‘same intent → one span’ rules determine what counts as coverage of ‘all eight dimensions’ and entry counts per product.
  • Temporal binning (2024 aggregate vs 2025 quarters; n=4 in 2026 dropped) = 2026 excluded (n=4)
    Binning and exclusion choices affect the reported growth narrative in Figure 5.
assumptions (5)
  • domain assumption Developer system prompts are a primary, persistent control layer that can override or reshape aligned base-model behavior in deployed products.
    Stated throughout §§1–2; underpins why prompt text audit is treated as user-protection evidence without paired behavioral A/B tests.
  • ad hoc to paper User-relevant prompt quality is adequately captured by eight UDHR-anchored dimensions with binary protective/problematic polarity on non-core-logic spans.
    §3 taxonomy design; alternative dimensions or graded severity would change coverage and prevalence statistics.
  • domain assumption Leaked or community-disclosed prompts, after maintainer and cross-repo checks, are valid objects for product- and organization-level claims.
    §5.1 and Appendix B; Limitations explicitly flags imperfect production fidelity.
  • domain assumption Presence of a protective instruction indicates meaningful adoption of that safeguard dimension at the prompt layer (without requiring proof of runtime enforcement).
    Used when reporting 98.9% ≥1 protective and per-dimension coverage in Figure 6.
  • standard math Standard qualitative reliability practices (calibration, IAA, expert adjudication) suffice to treat labels as stable enough for aggregate statistics.
    §5.1 reports pairwise IAA 0.933 on 20 spans; conventional content-analysis assumption.
invented entities (3)
  • AISPA eight-dimension span-level audit taxonomy
    purpose: Provide a unified protective/problematic coding scheme for system prompts tied to UDHR articles.
    Core methodological invention; dimensions remix known safety themes but the operational package is paper-specific.
  • Three-round LLM→annotator→expert prompt auditing protocol
    purpose: Scale span discovery while controlling false problematic labels.
    Procedural construct; validated only via internal IAA/adjudication, not external outcome benchmarks.
  • Gray-area / risky span category (four patterns)
    purpose: Hold borderline design trade-offs outside the ±1 main tallies.
    Post-hoc expert typology on 29 spans; useful descriptively but not independently measured against user harm.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AISPA: User-Centric System Prompt Auditing for Large Language Model Applications." pith.science (2026). https://pith.science/paper/YDN6QNZ3

@misc{pith2026260728617,
  author       = {Pith},
  title        = {Pith review of: AISPA: User-Centric System Prompt Auditing for Large Language Model Applications},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YDN6QNZ3}},
  note         = {Machine review of arXiv:2607.28617}
}
read the original abstract

System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.

Figures

Figures reproduced from arXiv: 2607.28617 by the authors.

Figure 1
Figure 1. Overview of system prompt quality across organizations. Bars show the average number of [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The eight auditing dimensions in AISPA. Each dimension targets a distinct aspect of responsible system prompt design. Our taxonomy comprises eight dimensions, each designed to capture both protective and problematic instructions within a single evaluative framework. For example, Identity Transparency includes both explicit disclosure that the system is AI (protective) and deliberate concealment of its AI nature (pro… view at source ↗
Figure 3
Figure 3. Illustration of span-level prompt auditing. A fictitious system prompt is segmented into spans. Each highlighted span is assigned a dimension and polarity under the taxonomy: the blue-toned span denotes a protective span (+1), while the red-toned span denotes a problematic span (−1). Unhighlighted sentences are not flagged by the auditor. 4.1 Auditing Guidelines Unit of analysis: Prompt Span. The basic unit of analy… view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Three-round collaborative audit protocol. In Round 1, an LLM pre-annotator identifies candidate spans with high recall. In Round 2, trained annotators screen these proposals to improve precision. In Round 3, three domain experts review the remaining cases to ensure con…
Figure 5
Figure 5. Figure 5: Temporal trends in system prompt evolution from 2024 to 2025. (a) Average number of +1 (protective) entries per product. (b) Average system prompt length in characters. (c) Percentage of products containing at least one −1 (problematic) entry. To ensure balanced repres…
Figure 6
Figure 6. Figure 6: Prevalence of user protection and problematic entries. (a) Product-level coverage: percent￾age of the 88 products that contain at least one +1 entry, at least one −1 entry, or +1 entries across all eight dimensions. (b) Dimension-level coverage: for each dimension, the…
Figure 7
Figure 7. Figure 7: Organization-level ranking based on the rating of system prompts. (a) Average protective spans per product. (b) Average problematic spans per product. Amazon and Cline follow with strong protective counts of 42.0 and 39.5, respectively, while having few problematic ins…
Figure 8
Figure 8. Figure 8: The evolution of system prompts for frontier models released by Anthropic, OpenAI and xAI. For Claude, GPT, and Grok model series, the system prompts gradually become more protective. Case study: version evolution across three providers [PITH_FULL_IMAGE:figures/full_f…
Figure 9
Figure 9. Figure 9: Distribution of 88 products across five categories. [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]
Figure 10
Figure 10. Figure 10: Cross-repository content overlap for each matched product. For each of the 88 prompts, we compute the Sørensen–Dice overlap with all same-product files in the other five repositories and retain the highest score. 50 prompts have at least one cross-repository match; 38…
Figure 11
Figure 11. Figure 11: Annotation platform example [PITH_FULL_IMAGE:figures/full_fig_p022_11.png]
Figure 12
Figure 12. Figure 12: Annotation platform example with human audited results. E Supplementary Analyses We report additional descriptive analyses that provide context for the main findings. These analyses characterize the structure of the audit dataset but do not constitute independent find…
Figure 13
Figure 13. Figure 13: Prompt size vs. user protection balance. Each bubble represents a product; size encodes the number of annotated spans and color encodes the protective rate. Products with the lowest protection balance are consistently those with short prompts. E.2 Dimension-Level Anal…
Figure 14
Figure 14. Figure 14: Word clouds for the eight auditing dimensions. Each panel shows the most frequent terms in annotated spans for that dimension, after removing domain-generic stopwords. Word Clouds show Dimension-Specific Vocabularies [PITH_FULL_IMAGE:figures/full_fig_p023_14.png]

Discussion (0). Sign in to comment.

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.