Selected references
Key works cited by the paper. Search, sort by column, or filter by year. Links open the source.
| Work | Authors | Year | Link |
|---|---|---|---|
| Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs | Betley et al. | 2025 | https://arxiv.org/abs/2502.17424 |
| Persona Features Control Emergent Misalignment | Wang et al. | 2025 | https://arxiv.org/abs/2506.19823 |
| The Persona Selection Model | Marks et al. | 2026 | https://alignment.anthropic.com/2026/psm/ |
| Natural Emergent Misalignment from Reward Hacking in Production RL | MacDiarmid et al. | 2025 | https://arxiv.org/abs/2511.18397 |
| Deliberative Alignment: Reasoning Enables Safer Language Models | Guan et al. | 2024 | https://arxiv.org/abs/2412.16339 |
| Constitutional AI: Harmlessness from AI Feedback | Bai et al. | 2022 | https://arxiv.org/abs/2212.08073 |
| Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training | Hubinger et al. | 2024 | https://arxiv.org/abs/2401.05566 |
| Monitoring Reasoning Models for Misbehavior | Baker et al. | 2025 | https://arxiv.org/abs/2503.11926 |
| HealthBench: Evaluating LLMs Towards Improved Human Health | Arora et al. | 2025 | https://arxiv.org/abs/2505.08775 |
| Inoculation Prompting: Train-Time Misbehavior Improves Test-Time Alignment | Wichers et al. | 2025 | https://arxiv.org/abs/2510.05024 |
| Stress Testing Deliberative Alignment for Anti-Scheming Training | Schoen et al. | 2025 | https://arxiv.org/abs/2509.15541 |
| DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors | Huang et al. | 2025 | https://arxiv.org/abs/2510.15501 |
| The MASK Benchmark: Disentangling Honesty From Accuracy in AI | Ren et al. | 2025 | https://arxiv.org/abs/2503.03750 |
| School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misaligned Behavior | Taylor et al. | 2025 | https://arxiv.org/abs/2508.17511 |
| PropensityBench: Evaluating Latent Safety Risks via an Agentic Approach | Sehwag et al. | 2025 | https://arxiv.org/abs/2511.20703 |
| Agentic Misalignment: How LLMs Could Be Insider Threats | Lynch et al. | 2025 | https://arxiv.org/abs/2510.05179 |
| Helpful Assistant Features Suppress Emergent Misalignment | Dupré la Tour | 2025 | https://alignment.openai.com/helpful-assistant-features/ |
| Sidestepping Evaluation Awareness with Production Evaluations | Williams et al. | 2025 | https://alignment.openai.com/prod-evals/ |
| Teaching Claude Why | Kutasov et al. | 2026 | https://alignment.anthropic.com/2026/teaching-claude-why/ |
| Positive Alignment: Artificial Intelligence for Human Flourishing | Laukkonen et al. | 2026 | https://arxiv.org/abs/2605.10310 |
| Concrete Problems in AI Safety | Amodei et al. | 2016 | https://arxiv.org/abs/1606.06565 |
| Truthful AI: Developing and governing AI that does not lie | Evans et al. | 2021 | https://arxiv.org/abs/2110.06674 |
| AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents | Andriushchenko et al. | 2024 | https://arxiv.org/abs/2410.09024 |
| GPQA: A Graduate-Level Google-Proof Q&A Benchmark | Rein et al. | 2023 | https://arxiv.org/abs/2311.12022 |
| SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Tasks? | Deng et al. | 2025 | https://arxiv.org/abs/2509.16941 |