Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:38:43.394556Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2607.26173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:38:43.394556Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fc3a5fa0-2556-45cf-88c1-7c0d305867e2 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models We give the generation procedure and matched examples here because the differences between conditions are the intervention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8676757c-90b4-4c7a-87db-da783520e8fa · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5dec608-6606-45fc-abbd-a41cec86ab9a · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a90eaa7-9c55-4b16-91ab-9fb2cfec81ef · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Safety Cases: How to Justify the Safety of Advanced AI Systems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f0fc32-df4b-4612-a852-66ed029f73a3 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762c5185-b624-4412-b4cc-bb2e2e16200c · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc22e798-f061-4c96-a7f9-41e756a7a177 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Scaling Laws for Autoregressive Generative Modeling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f585ea-d940-4e3c-9437-5ebcbc66a8e7 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6087d215-cef2-404e-ab66-eb634ae4b5eb · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44cd6584-eea2-4f5b-917b-3230b8ee86ba · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887ada73-7ec2-4d97-9977-7a02eb1f5c3a · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Andrew K
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fba030-3ade-47a3-9180-9bf59c3326ed · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Model Spec Midtraining: Improving How Alignment Training Generalizes
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542765ba-91c8-428e-a96c-5644e3497941 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Gradient Episodic Memory for Continual Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2615bce9-87ad-46dd-918d-5f8bcc90e66d · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a667ab4-e74f-4a76-a4b8-86779ba62199 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Negation Neglect: When models fail to learn negations in training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ff1159c-a17e-46fc-94a8-645f7070bc08 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Tell, don't show: Declarative facts influence how LLMs generalize
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f178cb15-eca4-4a32-8d2d-4c6a0401807b · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64369c4d-6cc6-430c-9667-54281e2bdc16 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models iCaRL: Incremental Classifier and Representation Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5ea07f-7d12-4c48-8ca3-f5cd63365fc3 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1885033-e6ae-4ec6-bb3f-25c27ca3ddaf · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e32811c-a2da-43b2-b25f-7939b4e2f58e · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Model Organisms for Emergent Misalignment
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb12c35-1561-450b-84e1-65c60c405e43 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab22fa55-c124-4552-8831-a7d159b6952d · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models LIMA: Less Is More for Alignment
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901e915a-a79d-4f43-9d52-e7f1fcc9008b · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d5d0e16-9c26-4354-9ca3-5472c066fb8a · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Experience Replay for Continual Learning
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e4bc95-2d42-4d3c-8011-73700833c500 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06789e2-a277-43d9-bd79-ce7a9ea4405e · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models On-Policy Replay for Continual Supervised Fine-Tuning
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d03c9a0-9501-4dad-bfbd-9608a5baab41 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Taken out of context: On measuring situational awareness in LLMs
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4adfa29e-c0b9-4ec2-b88d-aeeda5cc8705 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models BEiT: BERT Pre-Training of Image Transformers
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd9b9e6-fdb0-456c-b858-37eaf93e7919 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Constitutional AI: Harmlessness from AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18de718f-09dc-42c6-98f9-d936ada2e720 · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Looking Inward: Language Models Can Learn About Themselves by Introspection
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a7d092-57e3-4cd7-b651-c477d8a447be · outbound
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Scaling Laws for Generative Mixed-Modal Language Models
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.