Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:35.553049Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2502.08657.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:35.553049Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b90455ae-f6db-4334-8b9b-4dff47035811 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b33bec3d-afa7-4659-a17d-3d8fe19b393a · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions DeepSeek-V3 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44078d6b-9dfb-4df5-88da-a17a08d7b3e1 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions LLaMA: Open and Efficient Foundation Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07638cb-ec24-41fe-80bb-cdcd8e3a2928 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34533a29-e8b8-41ee-9f1e-e620a028d73b · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The ai alignment problem: why it is hard, and where to start,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0cc7659b-d836-47d5-9b15-275ffbaebbe6 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Artificial intelligence, values, and alignment,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c90bfd9-124d-4139-9ae0-dca01410a620 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27b1f92f-53ec-4980-af67-70b609d027eb · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-alignment of large language models via monopolylogue-based social scene simulation,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 648bc372-db72-460b-ba3d-ef72a90a9c22 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1901f4e-4d05-43f8-9993-53917d1a0576 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Balancing differential privacy and utility: A relevance-based adaptive private fine-tuning framework for language models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation efcc71b5-2986-463c-80ce-ffc8f2c53bc7 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Hard adversarial example mining for improving robust fairness,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 76e5cd17-b69e-40a1-b2d1-2e1afcbc0404 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Finetuned language models are zero-shot learners,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 805dd60b-66fe-4299-86d3-15c85574f5ed · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bde28a1-a72a-4264-bcc1-2f435ef56ae6 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-tuning aligned language models compromises safety, even when users do not intend to!
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39e23c0a-ca4b-435c-9abf-ce5dfc04f8f5 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Stanford alpaca: An instruction-following llama model,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34156006-5528-4174-af5e-f2ab8bc7a871 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00152d79-6299-4ce9-a8f6-718417e69234 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-instruct: Aligning language models with self- generated instructions,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2c4b06d-7662-4578-8b26-e836036081b2 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Training language models to follow instructions with human feedback,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8399b0d9-c134-414d-bce5-e7dddf93432e · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beavertails: Towards improved safety alignment of llm via a human-preference dataset,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1345c631-0cc0-4ecf-be07-6b52b9a6b7f5 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Poisoning language models during instruction tuning,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fdbce315-085b-4fea-be85-68c1556a2737 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openassistant conversations-democratizing large language model align- ment,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4b048045-507b-4655-9e95-31ed27700854 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Principle-driven self-alignment of language models from scratch with minimal human supervision,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d1dbda63-7d4d-4e17-a183-ab31d615a3df · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RAIN: Your Language Models Can Align Themselves without Finetuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0da6d42f-2839-4d31-8723-13b76ff2c4d6 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721bf31d-5100-4e67-830a-f891a999e833 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Self-concept clarity development across the lifespan,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8cec7ea0-ce83-4d76-89ee-807450d21061 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d604db-332f-4d2e-a985-bce45b67d65c · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Secrets of RLHF in Large Language Models Part II: Reward Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8c2ee2-4adb-404e-af5b-88f0a39a02f2 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Variational bayesian un- learning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f09253eb-a71a-4520-8631-7bcd66c226f4 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f879638-9652-4351-baca-4f6a1d159684 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions A Survey of Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92cdd58e-ad1c-4c04-8397-27b7704acb45 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions On the Opportunities and Risks of Foundation Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55eb9564-193f-4d61-aa02-4d7628961252 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9b95246-da35-47f7-a07a-4b72592d8596 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Fine-grained human feedback gives better rewards for language model training,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5f597f0-13b9-422e-87bb-785a3a9c9b4e · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning to summarize with human feedback,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 57659805-3e28-4c17-a855-6d5f70f22100 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 77cd4e56-c207-465f-b4b4-0d46f6b51d63 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gaining wisdom from setbacks: Aligning large language models via mistake analysis,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91041920-d5a2-45a2-8069-8d8b760dae0f · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Understanding negative samples in instance dis- criminative self-supervised representation learning,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4f2e0d2-7b2f-4c36-a474-aa7acc360373 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f9962e-41d9-4b82-8d17-eebc325c02ee · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions What makes for good views for contrastive learning?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ea4fda9-0edd-4acf-b225-23184ce5552e · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Neural text degeneration with unlikelihood training,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4cde7402-0787-42a0-8e7b-cf85aa83ef18 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Openai. gpt-4v(ision) system card,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b78a06e6-75c4-4d4a-9fb6-f9ec1518778e · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c835252-f39b-495d-b08e-1a973428ef69 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4bbf00f-7f71-48e1-8023-4e5969b47e7e · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Gemini: A Family of Highly Capable Multimodal Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9118a71-173d-4d90-94af-3407ca079fd2 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Chain-of-thought prompting elicits reasoning in large language models,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e369c3-388c-4625-afc6-ab1058cce931 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The wisdom of hindsight makes language models better instruction followers,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 340786a3-7012-4239-9582-dabbf100e490 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Koala: A dialogue model for academic research,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61f8cb44-e1ea-464a-900f-ba2eac9be164 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Glm-130b: An open bilingual pre-trained model,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4e281db-3fdd-4ff6-ada7-49ab27538497 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The curious case of neural text degeneration,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e9c7d08-786e-4f8e-9bac-9a6946ca5f81 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Lora: Low-rank adaptation of large language models,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ef50156-efc3-498c-a623-37771202d5ad · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions The Llama 3 Herd of Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fdce5e9-8d92-4dc6-9d1d-0bf8844b6519 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd711585-734e-43ee-8ba3-4e08b50fd056 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Autodan: Generating stealthy jailbreak prompts on aligned large language models,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf6c2c09-2ba9-4856-9d5f-4b273dcf9166 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf95be37-1915-4c0a-8207-3ac6d74ec4e3 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Truthfulqa: Measuring how models mimic human falsehoods,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 937b1e5a-7f3e-4358-b548-0f81431ab153 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Measuring massive multitask language understanding,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23b4cd2-24f3-4aec-a5ae-6609841d685b · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Social iqa: Commonsense reasoning about social interactions,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5534bcea-024b-408b-bde5-a85d0ea447b1 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Mitigating the alignment tax of rlhf,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 712610c3-5c13-4414-9edf-2fd41067a154 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions Direct preference optimization: Your language model is secretly a reward model,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bfb14e8-30c3-4297-bb0d-779945f7f404 · outbound
Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions KTO: Model Alignment as Prospect Theoretic Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.