{"id":"d2695b3e-f6d0-4135-8a4d-a7b69c3f0c24","arxiv_id":"2507.19185","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 198 community-collected prompts and nine commercial APIs, psychological manipulation outperformed technical obfuscation, attack transfer across models was limited to 16.9% of transformations, and Claude 4 models showed higher failure rates than earlier Claude versions.","lead":"The paper describes PrompTrend, a system that collects LLM jailbreak prompts from public forums and scores them with a new risk framework, then tests 198 collected prompts against nine commercial models. It reports that psychological manipulation beats technical tricks, that attacks rarely transfer between model families, and that in one model family newer versions look more vulnerable than older ones.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claude 'four-fold regression' is an unsupported temporal inference from a single cross-section of different model versions and tiers; removing it collapses the capability-security headline.","rationale":"I read the paper in good faith: the PrompTrend monitoring concept and the multi-dimensional scoring idea address a real gap, and the raw experiment may contain useful observations. The reader's weakest_assumption is also the most load-bearing concern I can identify. The conclusion's strongest claim—'capability advancement does not improve security,' concretized as a four-fold Claude regression—depends entirely on comparing different Claude versions and tiers collected in one snapshot, not on a measured change over time. The paper's own limitation section concedes the cross-sectional design, and the GPT-4o figure point is not traceable to the tested model set. Other concerns (PVAF thresholds recalibrated on the same data, execution-count inconsistencies, Eq. (1) omitting three of six advertised dimensions) are real and reinforce the reader's rejection, but the temporal inference is the single claim whose failure most directly collapses the central narrative. I do not see a need to change the reader's verdict: the paper should not be accepted in its current form. A revised version with longitudinal or repeated-measurement data, held-out PVAF validation, and a consistent model list could be reconsidered.","tokens_in":16316,"tokens_out":9969,"duration_ms":89851,"concrete_test":"Request or download the raw evaluation logs (the R and C matrices plus timestamps) from the GitHub repository and rebuild Figure 11b with x = actual execution date and separate markers for release date and model tier. Verify whether GPT-4o appears anywhere in the model set, then recompute the Claude failure-rate ratios after comparing only same-tier versions where available (e.g., Claude 3.5 Sonnet vs Claude 4 Sonnet) rather than Haiku vs Sonnet. If all nine models were tested in the same January-May 2025 window and GPT-4o is absent, the figure's temporal curve is a cross-sectional version comparison; the 'four-fold regression' sentence should be removed and the conclusion restated as a current cross-model difference, not a security trajectory.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 and Figure 11b present Claude Haiku (0.9%), Claude 3.5 Sonnet (1.3%), and Claude 4 Sonnet (4.1%) as a worsening security trajectory, and Section 6.1 and the Conclusion convert this into 'Claude 4's 4.1 percent vulnerability rate represents a four-fold regression from earlier versions.' The underlying data are a five-month cross-sectional snapshot (Section 4.1.1), and the paper itself concedes in Section 6.6 that 'the current study presents cross-sectional analysis.' No model version was measured repeatedly over time, and the compared models differ in tier (Haiku, Sonnet, Opus) as well as release date, so the apparent regression is confounded by model size/capability and version rather than being a measured change in security over time. The same figure's OpenAI trend cites GPT-4o at 1.9% in May 2024, but GPT-4o is not among the nine models listed in Section 4.2.1 (GPT-4 is), so the temporal curve includes a data point that cannot have been produced by this study. If the temporal framing is removed, the remaining result is a cross-sectional difference among current model versions; the headline claim 'capability advancement does not improve security' no longer follows.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PrompTrend, a monitoring system that collects LLM 'jailbreak' prompts from online communities (Reddit, Discord, GitHub, Twitter/X, security forums), filters them, and evaluates them with the PrompTrend Vulnerability Assessment Framework (PVAF). The authors report 198 unique vulnerabilities collected between January and May 2025, tested across nine commercial models using 71 transformation strategies. Their headline findings are that psychological manipulation outperforms technical obfuscation, that platform dynamics (especially Discord) shape attack effectiveness, that cross-model transferability is low, and that capability advancement does not improve security, with Claude 4 showing a 'four-fold regression' compared with earlier Claude versions. The paper also reports that PVAF achieves 78% classification accuracy, an AUC of 0.72, and a relative risk of 1.50 between low- and moderate-risk categories. The central contributions claimed are the monitoring architecture, the PVAF multidimensional scoring framework, and the empirical findings about model-specific vulnerability patterns.","tokens_in":16485,"tokens_out":4973,"duration_ms":45055,"significance":"If the central claims were sound, this would be a useful contribution: a community-driven vulnerability collection pipeline, a repeatable API testing protocol, and an attempt to incorporate social-adoption signals into vulnerability scoring are all valuable directions. The manuscript has visible strengths: Algorithm 1 gives a concrete testing protocol, the evaluation includes inter-rater reliability (Cohen's kappa = 0.76), the statistical reporting uses effect sizes and multiple-comparison corrections, and the paper includes ethical disclosure considerations. However, the two load-bearing claims do not survive scrutiny. The PVAF performance numbers come from a threshold choice made after seeing the data on the same dataset, making them circular; the 'capability advancement does not improve security' claim rests on treating a cross-sectional comparison of different model versions and tiers as a temporal trend, and Figure 11b even includes a model (GPT-4o) that is not in the tested set. The PVAF scoring definition is internally inconsistent, and several reported numbers do not reconcile. These problems reduce the significance of the empirical contributions as they stand.","major_comments":[{"comment":"The reported validation of PVAF is circular. Section 6.4 states that initial thresholds were recalibrated because vulnerabilities scoring 42–47 showed lower success rates than those scoring 34–41, and that the balanced terciles (0–33, 34–66, 67–100) 'restored monotonic risk progression'. The 78% accuracy, AUC 0.72, and relative risk 1.50 in Section 5.4 and Table 5 are computed on the same data that motivated the recalibration, with no held-out split, temporal replication, or independent test set. These numbers therefore do not establish predictive validity; they are in-sample goodness-of-fit statistics.","section":"§6.4 and §5.4"},{"comment":"The claim that Claude 4's 4.1% vulnerability rate is a 'four-fold regression' is an inference from a single cross-sectional snapshot, not a measured temporal trend. Section 4.1.1 and Section 6.6 describe the data as cross-sectional, and the three compared Claude models (Haiku, 3.5 Sonnet, 4 Sonnet) differ in tier and were each tested once; they are not repeated observations of the same system. Figure 11b also plots a GPT-4o data point dated May 2024, but Section 4.2.1 lists only GPT-4, O1, O3-Mini, and GPT-4.5 as the OpenAI models tested. Removing this unsupported temporal framing also removes the headline conclusion that capability advancement does not improve security.","section":"§5.1, Figure 11b, §6.1, Conclusion"},{"comment":"The PVAF scoring function is internally inconsistent and under-specified. Section 3.3.1 defines six dimensions with weights 0.20, 0.20, 0.15, 0.15, 0.15, 0.15, and Section 3.3.2 adds dynamic modifiers, but Eq. (1) in Section 4.3 is a three-term sum with equal weights 0.33 and omits Cross-Platform Efficacy, Temporal Resilience, Propagation Velocity, and all dynamic modifiers. The manuscript never states how the six dimensions are aggregated, how Community Adoption is computed from engagement signals, or whether dynamic modifiers were applied in the reported scores. In addition, the engagement-based Community Adoption dimension is derived from the same filtering signals used to select the 198 prompts (Sections 3.2.1 and 4.1.2), introducing circularity into the scoring. The reported PVAF values are therefore not reproducible from the description.","section":"§3.3.1, §3.3.2, §4.3, Eq. (1)"},{"comment":"Several numerical results do not reconcile with each other. Table 5 reports a moderate-risk success rate of 16.90% and low-risk 11.27% on 22,152 test executions, while Section 5.3 gives an overall mean success rate of 2.0%; Section 4.2.2 reports 199,368 total executions. The text does not explain why the Table 5 denominator is 22,152, whether the success-rate definitions are the same, or how the 2.0% overall rate can coexist with the Table 5 rates. The high-risk row of Table 5 simultaneously reports '—' for the success rate and 0% of tests, which is not a meaningful comparison. These inconsistencies make the headline PVAF stratification result difficult to verify.","section":"§5.4, Table 5, §4.2.2, §5.3"}],"minor_comments":[{"comment":"Section 4.1.2 reports 2,800 unique vulnerabilities qualifying for detailed PVAF assessment, while Section 4.1.1 reports 198 unique vulnerability prompts; clarify the relationship between these numbers.","section":"§4.1.2 and §4.1.1"},{"comment":"Figure 11b is labeled as temporal evolution, but the underlying data are cross-sectional; the figure should be replotted as a model-version comparison or removed.","section":"Figure 11b"},{"comment":"Section 3.2.2 states that bridge-node identification is 'designed for...future deployments' while the surrounding text describes it as part of the implemented system; align the tense and scope.","section":"§3.2.2"},{"comment":"In Table 5, the High Risk row should use 'not observed' rather than '—' for the success rate, and the 0% test count should be explained.","section":"Table 5"},{"comment":"The statement in Section 6.5 that 'Claude 4 regresses' repeats the invalid temporal interpretation identified in Major Comment 2 even after the Limitations section.","section":"§6.5"}],"recommendation":"reject","confidential_remarks":"The two load-bearing claims (PVAF predictive validity and the capability-security regression) rest on circular or cross-sectional evidence as detailed in the major comments. A revision would require fresh validation with a held-out set and a fundamental reframing of the temporal claims; as it stands the manuscript is not suitable for publication. I would not ask the authors for minor edits only."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look for the concept, not for the empirical claims. The idea of continuously scraping Reddit, Discord, GitHub, and Twitter for jailbreak prompts and scoring them along social-technical dimensions is sensible and underexplored. The system description is detailed, and the transformation taxonomy is reasonably comprehensive. That is the real contribution, and it is a contribution of design, not of evidence.\n\nThe reader's take is largely correct. The validation is circular: Section 6.4 says thresholds were recalibrated after seeing that 42–47 scored prompts underperformed 34–41, and the 78% accuracy, AUC 0.72, and relative risk are all reported on the same data. That is not prediction; it's curve fitting. Equation (1) defines PVAF with only three dimensions (HP, ES, CA) and equal weights 0.33, but the text advertises six dimensions and different weights. Either the equation or the prose is wrong, and that matters because the framework's credibility rests on the score.\n\nThe temporal claims are the weakest part. Figure 11b and the conclusion say Claude 4's 4.1% is a 'four-fold regression' from earlier versions, but the data are a five-month cross-section of different model tiers (Haiku, Sonnet, Opus). No model was measured over time. Worse, the same figure cites GPT-4o at 1.9% in May 2024, a model that is not in the test set listed in Section 4.2.1. That point cannot have come from this study, and no source is given. Without the temporal framing, the 'capability advancement does not improve security' headline collapses into a static comparison of current models.\n\nThere are also numeric inconsistencies: the theoretical test space is 126,414 (198×71×9), but the paper reports 199,368 actual executions and later 22,152 test executions for the PVAF analysis. Those numbers cannot all be right. The dataset link is a GitHub repository, but no data or code is released in usable form, and the 'first dataset' claim is not benchmarked against Shen et al.'s in-the-wild jailbreak dataset, which the paper itself cites.\n\nWhat holds up? The cross-sectional result that psychological techniques outperform technical obfuscation is consistent with GUARD and CyberArk, so the direction is credible even if the current study can't validate it. Limited cross-model transfer (16.9%) is also plausible and matches Kim et al. The paper is honest in Section 6.6 that the study is cross-sectional, which is good, but then the conclusion overstates it anyway.\n\nMy verdict: reject in current form, but not a desk-reject. The concept deserves a serious referee, and the authors could come back with a held-out validation set, corrected equations and counts, released artifacts, and a re-framed cross-sectional analysis. As it stands, the load-bearing claims are fitted and the temporal inference is unsupported.","headline":"A useful monitoring concept undermined by circular validation and unsupported temporal claims.","tokens_in":17121,"tokens_out":2326,"would_cite":false,"duration_ms":21842,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the dominant LLM threat is psychological manipulation from online communities, that successful attacks rarely transfer across model families, and that, within the Claude family, newer models are more vulnerable than…","keywords":["LLM security","vulnerability assessment","community-driven discovery","jailbreak attacks","psychological manipulation","continuous threat intelligence","AI safety","risk scoring"],"falsifier":"Rerun the same 198-prompt corpus and 71 transformations against an older and a newer model of the same family under identical API settings; if the newer model's failure rate is not higher, the claimed four-fold regression collapses.","tokens_in":16010,"feed_emoji":"🎭","tokens_out":8949,"duration_ms":79297,"temperature":0.7,"pith_summary":"This paper tries to establish that the real threat to commercial large language models is being discovered in online communities, not in formal benchmarks, and that a continuous monitoring system can capture it. The authors built PrompTrend, which collected 198 vulnerability prompts from chat, forum, and code-sharing platforms over five months and tested them across nine models using 71 transformation strategies. Their central finding is that psychological manipulation, such as emotional appeals and roleplay, succeeds at roughly twice the rate of technical obfuscation like encoding or prefix injection. They also report that successful attacks rarely transfer across model families, and that among the Claude models tested the newest versions failed more often than earlier ones. The paper matters because it challenges the assumption that newer, more capable models are automatically safer.","feed_headline":"Emotional prompts break LLMs twice as often as technical ones","feed_subtitle":"Community-sourced attacks show newer Claude models are more vulnerable and attacks rarely transfer between model families.","key_machinery":"The central mechanism is the PrompTrend pipeline paired with the PVAF scoring framework. PrompTrend is a three-stage collection, enrichment, and scoring system whose platform-specific agents filter social-media streams for vulnerability candidates, producing structured metadata for each candidate. PVAF evaluates each vulnerability across six dimensions, including harm potential, exploit sophistication, community adoption, cross-platform efficacy, temporal resilience, and propagation velocity, and combines them into a 0-100 risk score. The framework's distinguishing work is adding social dynamics to technical risk, and the paper validates it by showing that higher PVAF scores predict higher measured jailbreak success, with a relative risk of 1.5 between moderate- and low-risk categories and an AUC of 0.72.","core_discovery":"The paper's core claim is that capability advancement does not guarantee security improvement and that community-driven psychological manipulation is the dominant vulnerability vector for current LLMs. Using 198 community-discovered vulnerabilities tested with 71 transformations on nine commercial models, the authors report a vulnerability hierarchy with Claude 4 Sonnet highest at 4.1% and GPT-4.5 lowest at 0.6%, with emotional manipulation at 4.9% success versus Base64 at 2.7%. They also report that chat-platform-sourced attacks are most effective against Claude-family models, that only 16.9% of successful attacks transfer across model families, and that a multidimensional scoring framework called PVAF reaches 78% classification accuracy. The authors argue these results demonstrate that static benchmarks miss the evolving, socially driven threat landscape and that security evaluation must be continuous and community-aware.","pith_inferences":["My inference: the claimed regression of the Claude family is a cross-sectional comparison of different current model versions, so a longitudinal release-by-release test is needed before treating newer-is-less-safe as a general law.","My inference: the emotional-manipulation result could be tested against a benign-emotion control; if emotionally framed benign requests also raise compliance, the mechanism is over-helpfulness rather than a safety-policy gap.","My inference: the platform-channel effects, with chat-borne psychological attacks outperforming code-repository technical attacks, imply that monitoring priorities could be tuned per deployment, though the small effect sizes caution against large operational bets.","My inference: the 16.9% cross-model transferability figure may undercount universal attacks because the transformation taxonomy was partly seeded by the same community patterns being measured."],"forward_implications":["If psychological attacks dominate, red-team exercises and safety training should be rebuilt around social-engineering scenarios rather than token-level perturbations.","If successful attacks rarely transfer across model families, defenses must be model-specific and universal jailbreak defenses are unlikely to suffice.","If the capability-security inversion within the Claude family is real, deployment decisions should not treat newer versions as automatically safer.","If PVAF risk scores predict jailbreak success, security teams can triage community-discovered prompts by score before full empirical testing.","If static benchmarks miss community-driven discovery, continuous platform monitoring becomes a necessary complement to formal evaluation."],"supporting_citations":[{"why":"supplies the in-the-wild jailbreak corpus and the DAN example that grounds the claim that community experimentation precedes formal analysis","marker":"[38]"},{"why":"defines the static red-teaming benchmark whose snapshot limitation PrompTrend is designed to overcome","marker":"[24]"},{"why":"provides the automated red-teaming baseline contrasted with community-driven discovery","marker":"[32]"},{"why":"gives the gradient-based attack method whose unnatural prompts stand in for technical obfuscation in the comparison","marker":"[49]"},{"why":"shows role-play generators rank persuasion-based jailbreaks above token-level perturbations, supporting psychological dominance","marker":"[19]"},{"why":"documents a production-chatbot exploit of emotional framing against deployed assistants","marker":"[40]"},{"why":"supplies the behavioral-science basis for treating social dynamics as a vulnerability dimension","marker":"[33]"},{"why":"provides the prior cross-platform transferability result that the paper's 16.9% transfer finding extends","marker":"[22]"},{"why":"explains the many-shot jailbreak mechanism used to attribute higher Claude vulnerability to a larger attack surface","marker":"[15]"}],"fun_headline_variants":["Emotional prompts dominate LLM attack success","Newer Claude models more vulnerable to community attacks","Capability gains don't shield LLMs from psychological attacks","Most LLM attacks fail to transfer across model families","Static benchmarks can't keep pace with community-driven LLM threats"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's central warning depends on treating a single snapshot of different current model versions as a trend over time, not on repeated measurements of the same system.","fun_headline_variants_meta":{"raw":{"variants":["Emotional prompts dominate LLM attack success","Newer Claude models more vulnerable to community attacks","Capability gains don't shield LLMs from psychological attacks","Most LLM attacks fail to transfer across model families","Static benchmarks can't keep pace with community-driven LLM threats"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1412,"prompt_tokens":853,"completion_tokens":559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":483}},"tokens_in":469,"tokens_out":559,"duration_ms":6186,"temperature":1.0,"reasoning_tokens":483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:59:39.206754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same 198-prompt corpus and 71 transformations against an older and a newer model of the same family under identical API settings; if the newer model's failure rate is not higher, the claimed four-fold regression collapses.","supporting_citations":[{"cited_title":"HarmBench: Astandardizedevaluationframeworkforautomatedredteaming and robust refusal","cited_arxiv_id":null,"evidence_quote":"defines the static red-teaming benchmark whose snapshot limitation PrompTrend is designed to overcome"},{"cited_title":"Shimony and S","cited_arxiv_id":null,"evidence_quote":"documents a production-chatbot exploit of emotional framing against deployed assistants"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the behavioral-science basis for treating social dynamics as a vulnerability dimension"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"explains the many-shot jailbreak mechanism used to attribute higher Claude vulnerability to a larger attack surface"}],"review_version":2}