REVIEW 21 cited by
International AI Safety Report
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by the nations attending the AI Safety Summit in Bletchley, UK. Thirty nations, the UN, the OECD, and the EU each nominated a representative to the report's Expert Advisory Panel. A total of 100 AI experts contributed, representing diverse perspectives and disciplines. Led by the report's Chair, these independent experts collectively had full discretion over the report's content.
Forward citations
Cited by 21 Pith papers
-
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
-
SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts
Frontier LLMs pass fewer than 58% of systematically varied safety-fact scenarios, revealing weak generalization of critical safety knowledge to naive user queries.
-
Hardware Mechanisms to Dynamically Throttle AI Performance
Dynamic microarchitecture throttling of GPU memory resources can cut LLM inference performance by up to 80% with low hardware overhead, giving architects a continuous, hardware-enforced AI capability control.
-
Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.
-
Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis
A game-theoretic model shows media can act as a soft regulator of AI safety, but only when media signals are reliable and costs are low; otherwise defection can persist.
-
The Other Mind: How Language Models Exhibit Human Temporal Cognition
Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.
-
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
On Llama-3 and Qwen-2.5, removing safety guardrails sharply raises compliance with dangerous bio, chem, and cyber requests, and the resulting safety gap grows with model scale.
-
Technical Options for Flexible Hardware-Enabled Guarantees
A hardware 'interlock' placed on AI accelerator network paths could provide privacy-preserving, verifiable guarantees about AI compute usage, according to a design analysis that sketches FLOP-counting and update protocols.
-
Reconsidering LLM Uncertainty Estimation Methods in the Wild
Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.
-
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Attacks that break LLMs best are not the ones that improve safety most; a Shapley- and greedy-based framework that selects attack subsets by downstream defender utility outperforms attacker-centric and attribution-onl...
-
On the Principles of Deep Feedforward ReLU Networks
Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.
-
The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
AI research agents remove the human accountability backstop that prior science verification relied on, so the paper proposes observable-by-default workflows, tiered verification, and AI attribution standards to preser...
-
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
Across 23 models from 135M to 32B parameters, internal factual knowledge scales about twice as fast with model size as linguistic competence, supporting modular small-model-plus-retrieval systems.
-
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.
-
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.
-
A Taxonomy of Omnicidal Futures Involving Artificial Intelligence
A conceptual taxonomy dividing AI-driven omnicide into unintentional, intentional-by-state, intentional-by-institution, intentional-by-individual, and intentional-by-AI scenarios.
-
LLM Agents Should Employ Security Principles
A position paper proposing AgentSandbox, a framework that applies Saltzer-Schroeder security principles to LLM agents and reports large attack-success-rate reductions on AgentDojo.
-
Mitigating Deceptive Alignment via Self-Monitoring
CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.
-
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI
AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.
-
Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey
A survey mapping how neurosymbolic AI could address safety, regulatory, and operational challenges in advanced air mobility.
-
From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs
LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.
Discussion (0). Sign in to comment.