Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.