Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:13:07.507975Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 1 inbound Pith citation observation for arXiv:2412.14186.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T20:13:07.507975Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T16:21:29.283270Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 102 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ab8e546e-e28e-4db1-a330-e76a9a2fc606 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://www.anthropic.com/news/anthropics-responsible-scaling- policy, 2023
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a2263c-de56-42e9-87ed-9ead3dc8aa32 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://idais.ai/dialogue/idais- beijing/, 2024
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58a51031-9896-4d7e-b7fd-35b6c2fd46ed · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://idais.ai/dialogue/idais-venice/, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbdeecb-1309-4a0c-a253-0c0c4e33fdec · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI https://assets.anthropic.com/m/24a47b00f10301cd/original/Anthropic- Responsible-Scaling-Policy-2024-10-15.pdf, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 281d96d5-9922-4d53-a73a-8a28f39c054e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Current state of LLM Risks and AI Guardrails
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bd7614-50ff-4f65-a2e4-d3031ca986b5 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b5b63e3-ae89-405e-a6ee-2628632694fa · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d221283d-5433-4938-852d-3fecd7b36ac9 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Constitutional AI: Harmlessness from AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e850713-c1f6-4d4d-b0eb-b24d1e45cd17 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Managing extreme ai risks amid rapid progress
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889594c5-99e1-4e59-9fdc-cd5782c82624 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mechanistic Interpretability for AI Safety -- A Review
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63d28da4-a34e-4318-afa3-1f03d83a820b · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Diverse and effective red teaming with auto-generated rewards and multi-step reinforcement learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236f959f-c4e5-4f57-8810-375533c6a15a · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Measuring Progress on Scalable Oversight for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdbfbfe4-37dc-4e76-ac13-77a5baf7e6d3 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Language Models are Few-Shot Learners
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e58e1d-7919-4a8f-9c68-74d2bba1f73d · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262f3c42-fa6c-49ab-a53a-dad64ce2a799 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b83562e4-d564-4e7d-816c-bc01df2926fa · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI InternLM2 Technical Report
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a3d488-2459-4103-bab2-60d6ee1f6309 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Extracting training data from large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ee7580-0ebe-4f16-8a2b-135b98a23351 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Is Power-Seeking AI an Existential Risk?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f759c302-4a0f-4278-a65f-33b9d4d1fef2 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60224cab-b03b-4c4d-96d3-6579563faebe · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Cello: Causal evaluation of large vision- language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6b7c5d-e106-480c-8717-670a368b630e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI From Imitation to Introspection: Probing Self-Consciousness in Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629e536e-6ea7-4671-beca-a0e13956bd71 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b229e7d4-f9f5-411b-853e-31bbf6c9f522 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Feder Cooper, Christopher A
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b76f71-7b68-47b2-b794-bb5d58aa89d8 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1496d89d-e536-41f9-99df-c3d4439a4620 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Arti- ficial intelligence regulation: a framework for governance
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b97befef-d27a-4dd4-91b3-a84e2eacb63d · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c68b53-5315-4ccd-bdb0-d1e3ded9c867 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96721cf1-8b99-466c-9273-56eeb7809bef · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c775fd-b12e-4977-90d1-3386ab554974 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI CLEAR: Character Unlearning in Textual and Visual Modalities
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5821417-8506-4f86-916a-e7ae49a6dc5a · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Llama 3 Herd of Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ab911b-f204-43f1-b53c-a0627c33c8fe · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Who's Harry Potter? Approximate Unlearning in LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34283bd-422d-4cda-9ea6-f68ba01fcd37 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Statement on ai risk, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a212e10-a178-46d7-9d30-e1f3253a6bc1 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Counterintuitive behavior of social systems
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f78bdb-adc2-4a7c-9db6-765f087e1b8c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Artificial intelligence, values, and alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44781c90-f85b-42dc-bbc7-0255d4a181d8 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mental models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c6463b0-eb65-496e-b44b-07a88f24e229 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mllmguard: A multi-dimensional safety evaluation suite for multimodal large language models, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e355bb79-0c2d-48c1-93f1-45c9bcbf154d · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Recurrent world models facilitate policy evolution
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4c07748-652f-45fb-98d3-c01ae984de9b · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI An Overview of Catastrophic AI Risks
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54134020-7c5e-4841-8da0-db7ab7cff9b9 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Stabilizing translucencies: Governing ai transparency by standardization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2d6cd27-a09b-40b7-ab78-1cee5c703f99 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Curiosity-driven Red-teaming for Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80aa5f79-19da-4647-bd0e-e171fa565044 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Flames: Benchmarking Value Alignment of LLMs in Chinese
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89137e52-1153-4bc3-86f2-d01a364fb5b4 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI From pixels to principles: A decade of progress and landscape in trustworthy computer vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8523b380-a40f-4292-8cba-5ff8b12a510b · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI TrustLLM: Trustworthiness in Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edb1d758-75ab-4852-9981-5c1c3e923e19 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI When code isn’t law: rethinking regulation for artificial intelligence
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283c55fa-0836-4387-b2f3-17ca69c13622 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Scaling Laws for Neural Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f54af0a-9567-42a7-a19a-72673ac5eb22 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Aligning Large Language Models with Representation Editing: A Control Perspective
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5529e7-ddff-4b71-b5e3-91683dd7df83 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Evaluations: autonomy and artificial intelligence: a threat or savior? Springer, 2017
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8444ec68-984f-4367-8dab-ce953c9145a2 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Learning to watermark llm-generated text via rein- forcement learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 00fdb958-fd79-4e4b-acbc-1649665b3082 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Deepfakes, phrenology, surveillance, and more! a taxonomy of ai privacy risks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e68eff2c-9636-462c-9957-a9c9d94546de · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Trustworthy ai: From principles to practices
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7e39809-8f23-4808-b418-51f7bb2a7b08 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Inference- time intervention: Eliciting truthful answers from a language model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ba8dac-baab-4d56-93c4-e2db98710fb7 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b155ab8-2e38-4823-8f44-c8995285353c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Controllable Text Generation for Large Language Models: A Survey
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07a52096-c47a-4aad-b372-5ec55040a17f · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI A survey of text watermarking in the era of large language models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7eb4a8b-0692-4ed1-b7fe-0adecf89127b · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Don’t always say no to me: Benchmarking safety-related refusal in large vlm
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1271534e-75c2-479f-abdd-6aff9529ade7 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Mm-safetybench: A benchmark for safety evaluation of multimodal large language models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65490ca-8a6c-4ca7-9769-7301d70a9cfd · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689e8dc2-d173-4654-887a-6c7d80f166a7 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05e04c1e-0eae-4277-8747-463fc9acd83c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Machine Unlearning in Generative AI: A Survey
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e6973a-ccf5-44ce-9459-5c0cec0f506e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Inference-Time Language Model Alignment via Integrated Value Guidance
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f31435b5-8348-4a73-8175-d677027295e4 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Causal interpretability for machine learning-problems, methods and evaluation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55a5e382-a4d6-405a-a976-34a5494ab796 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Large Language Models in Cybersecurity: State-of-the-Art
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be1af07-4abb-4b01-a920-0ab8dbe4f85d · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Rule Based Rewards for Language Model Safety
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524aa9ec-29a2-45c5-a1c0-3a4c0dd07805 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Accountability in artificial intelligence: what it is and how it works
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 195be6b6-8310-443f-ae8f-c415c63e20a0 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676e3913-809f-4b4b-9c04-dc6baecc17a9 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI GPT-4 technical report
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c75a379a-39c4-4b3f-921f-36f0185ce1e3 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Openai o1 system card, 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3837409-4563-4be1-a293-d69f35d01da3 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Video generation models as world simulators, 2024
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 17ad1dea-471c-480c-bb83-b909736bb710 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Training language models to follow instructions with human feedback
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60f9a0a0-f821-48ab-9bf8-86c3c2de378c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI A ‘biased’emerging governance regime for artificial intelligence? how ai ethics get skewed moving from principles to practices
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3aed9a6-5a73-4d0f-83ed-d4ab18aae8be · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19c3bd3-579c-4dab-a6c1-ee7d5237ce5c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16bb05c-2e1e-434b-85ec-fc4e2aa33710 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Causality
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f76e6e-35a8-417c-97d3-025aaf4dc9f2 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The book of why: the new science of cause and effect
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769713d2-1ffc-4ad1-acfa-1bd846349a73 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Red Teaming Language Models with Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e9f853-911d-49f4-b09d-55b6d9e1dc19 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Tug of War Within: Mitigating the Fairness-Privacy Conflicts in Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d15e6ae2-5b79-456d-a11f-6de72e10181c · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49771c48-993c-4447-9870-c923c6743ed1 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Direct preference optimization: Your language model is secretly a reward model
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151fcb9d-a2b4-4a70-aac3-00df20dfe079 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Natural language processing: transform- ing how machines understand human language (2023)
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca19857c-496d-4e75-8bdd-b225bedc72f2 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Identifying Semantic Induction Heads to Understand In-Context Learning
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a9bea55-8d9c-4049-b80a-02e9a47ebd29 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Self-Reflection in LLM Agents: Effects on Problem-Solving Performance
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6680b5c9-4332-47e4-902b-d10b72e4bf9e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Scaling Laws for Deep Learning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3efea1a6-9bdf-44c0-941e-e8f650ae0132 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Re- flexion: Language agents with verbal reinforcement learning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34d9f7ab-1a35-413c-a767-d6663dfcc92e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Learning to summarize with human feedback
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e56d77-ca8f-4a3e-9fd0-98470ed9e8fb · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Reinforcement learning: An introduction
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6c44c11f-f240-4c30-97ce-93b98df58aeb · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI CAT-LLM: Style-enhanced Large Language Models with Text Style Definition for Chinese Article-style Transfer
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f9e657-7e39-4324-8e44-f4fc91b14386 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc986fd7-fa04-4f3f-82bf-ac25556e8fea · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Value sensitive design and responsible innovation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35940662-954e-4a70-849f-92b532d02452 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Interpretable counterfactual explanations guided by prototypes
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02a5e747-6ba4-4c50-939a-b2944992cd86 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9ca466-f652-46ac-9542-edbe4183b4a4 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Decodingtrust: A comprehensive assess- ment of trustworthiness in gpt models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1d89373d-94c2-434b-af8d-6d18ed66cf22 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318432f9-31a9-43b0-a95a-fa2b0538aa85 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Emu3: Next-Token Prediction is All You Need
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44223556-8486-4980-a048-5f8035dd494e · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI ai safety as global public goods
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7337fddf-e43e-412b-aca7-f4f38d81bc88 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Using the veil of ignorance to align ai systems with principles of justice
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44874b7e-63ac-4ab6-acb2-0fa4c39cdb27 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b9adf3-ae9b-40a8-a187-bc28328e9873 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Qwen2 Technical Report
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bcf0b3-c09a-4a53-a88e-4b5bb5af8f4a · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI Huref: Human-readable fingerprint for large language models
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f62cc241-170d-4d78-bebe-41903a85de1f · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI The Better Angels of Machine Personality: How Personality Relates to LLM Safety
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed64c77-9407-45b1-9320-5ba6919573f3 · outbound
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI REEF: Representation Encoding Fingerprints for Large Language Models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b10ab4f7-b386-4150-b888-57121a19ab40 · inbound
An Early Warning of Emerging Biosecurity Risks in Frontier LLMs Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.