Pith. sign in

REVIEW 21 cited by

International AI Safety Report

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.17805 v1 pith:VVBXBU7O submitted 2025-01-29 cs.CY cs.AIcs.LG

Yoshua Bengio , Sören Mindermann , Daniel Privitera , Tamay Besiroglu , Rishi Bommasani , Stephen Casper , Yejin Choi , Philip Fox
show 88 more authors
This is my paper
classification cs.CYcs.AIcs.LG
keywords reportsafetyexpertsinternationalnationsadvancedadvisoryattending
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The first International AI Safety Report comprehensively synthesizes the current evidence on the capabilities, risks, and safety of advanced AI systems. The report was mandated by the nations attending the AI Safety Summit in Bletchley, UK. Thirty nations, the UN, the OECD, and the EU each nominated a representative to the report's Expert Advisory Panel. A total of 100 AI experts contributed, representing diverse perspectives and disciplines. Led by the report's Chair, these independent experts collectively had full discretion over the report's content.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

    cs.CL 2026-07 accept novelty 7.0 of 10

    Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.

  2. SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts

    cs.AI 2025-05 conditional novelty 7.0 of 10

    Frontier LLMs pass fewer than 58% of systematically varied safety-fact scenarios, revealing weak generalization of critical safety knowledge to naive user queries.

  3. Hardware Mechanisms to Dynamically Throttle AI Performance

    cs.AR 2026-07 conditional novelty 6.0 of 10

    Dynamic microarchitecture throttling of GPU memory resources can cut LLM inference performance by up to 80% with low hardware overhead, giving architects a continuous, hardware-enforced AI capability control.

  4. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

  5. Can Media Act as a Soft Regulator of Safe AI Development? A Game Theoretical Analysis

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A game-theoretic model shows media can act as a soft regulator of AI safety, but only when media signals are reliable and costs are low; otherwise defection can persist.

  6. The Other Mind: How Language Models Exhibit Human Temporal Cognition

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Larger LLMs develop a subjective 'present' around the current date, and their year similarity judgments follow a logarithmic Weber-Fechner compression, with supporting neural and representational evidence.

  7. The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

    cs.CY 2025-07 conditional novelty 6.0 of 10

    On Llama-3 and Qwen-2.5, removing safety guardrails sharply raises compliance with dangerous bio, chem, and cyber requests, and the resulting safety gap grows with model scale.

  8. Technical Options for Flexible Hardware-Enabled Guarantees

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A hardware 'interlock' placed on AI accelerator network paths could provide privacy-preserving, verifiable guarantees about AI compute usage, according to a design analysis that sketches FLOP-counting and update protocols.

  9. Reconsidering LLM Uncertainty Estimation Methods in the Wild

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.

  10. How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

    cs.CR 2026-07 conditional novelty 5.0 of 10

    Attacks that break LLMs best are not the ones that improve safety most; a Shapley- and greedy-based framework that selects attack subsets by downstream defender utility outperforms attacker-centric and attribution-onl...

  11. On the Principles of Deep Feedforward ReLU Networks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Deep feedforward ReLU networks generalize two-layer principles via paths, piecewise linear manifolds, and continuity restriction to explain training solutions.

  12. The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

    cs.CY 2026-06 conditional novelty 5.0 of 10

    AI research agents remove the human accountability backstop that prior science verification relied on, so the paper proposes observable-by-default workflows, tiered verification, and AI attribution standards to preser...

  13. Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?

    cs.CL 2025-09 reject novelty 5.0 of 10

    Across 23 models from 135M to 32B parameters, internal factual knowledge scales about twice as fast with model size as linguistic competence, supporting modular small-model-plus-retrieval systems.

  14. TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Trident-Bench provides 2,652 professionally validated harmful prompts across finance, law, and medicine, and shows that domain-specialized LLMs often comply with unethical requests more than generalist models.

  15. Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

    cs.AI 2025-07 conditional novelty 5.0 of 10

    An evaluation of 18 frontier AI models across seven catastrophic-risk categories finds all models in green or yellow zones, with none crossing the report's proposed red lines.

  16. A Taxonomy of Omnicidal Futures Involving Artificial Intelligence

    cs.AI 2025-07 accept novelty 5.0 of 10

    A conceptual taxonomy dividing AI-driven omnicide into unintentional, intentional-by-state, intentional-by-institution, intentional-by-individual, and intentional-by-AI scenarios.

  17. LLM Agents Should Employ Security Principles

    cs.CR 2025-05 conditional novelty 5.0 of 10

    A position paper proposing AgentSandbox, a framework that applies Saltzer-Schroeder security principles to LLM agents and reports large attack-success-rate reductions on AgentDojo.

  18. Mitigating Deceptive Alignment via Self-Monitoring

    cs.AI 2025-05 conditional novelty 5.0 of 10

    CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.

  19. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0 of 10

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

  20. Integrating Neurosymbolic AI in Advanced Air Mobility: A Comprehensive Survey

    cs.RO 2025-08 conditional novelty 4.0 of 10

    A survey mapping how neurosymbolic AI could address safety, regulatory, and operational challenges in advanced air mobility.

  21. From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs

    cs.CR 2025-06 conditional novelty 3.0 of 10

    LLMs can assist both attackers and defenders in cybersecurity, but context limits, hallucinations, and weak reasoning make them unsafe to deploy without human oversight and real-world evaluation.

Pith tools