Selected references

Key works cited by the paper. Search, sort by column, or filter by year. Links open the source.

WorkAuthorsYearLink
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsBetley et al.2025https://arxiv.org/abs/2502.17424
Persona Features Control Emergent MisalignmentWang et al.2025https://arxiv.org/abs/2506.19823
The Persona Selection ModelMarks et al.2026https://alignment.anthropic.com/2026/psm/
Natural Emergent Misalignment from Reward Hacking in Production RLMacDiarmid et al.2025https://arxiv.org/abs/2511.18397
Deliberative Alignment: Reasoning Enables Safer Language ModelsGuan et al.2024https://arxiv.org/abs/2412.16339
Constitutional AI: Harmlessness from AI FeedbackBai et al.2022https://arxiv.org/abs/2212.08073
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingHubinger et al.2024https://arxiv.org/abs/2401.05566
Monitoring Reasoning Models for MisbehaviorBaker et al.2025https://arxiv.org/abs/2503.11926
HealthBench: Evaluating LLMs Towards Improved Human HealthArora et al.2025https://arxiv.org/abs/2505.08775
Inoculation Prompting: Train-Time Misbehavior Improves Test-Time AlignmentWichers et al.2025https://arxiv.org/abs/2510.05024
Stress Testing Deliberative Alignment for Anti-Scheming TrainingSchoen et al.2025https://arxiv.org/abs/2509.15541
DeceptionBench: A Comprehensive Benchmark for AI Deception BehaviorsHuang et al.2025https://arxiv.org/abs/2510.15501
The MASK Benchmark: Disentangling Honesty From Accuracy in AIRen et al.2025https://arxiv.org/abs/2503.03750
School of Reward Hacks: Hacking Harmless Tasks Generalizes to Misaligned BehaviorTaylor et al.2025https://arxiv.org/abs/2508.17511
PropensityBench: Evaluating Latent Safety Risks via an Agentic ApproachSehwag et al.2025https://arxiv.org/abs/2511.20703
Agentic Misalignment: How LLMs Could Be Insider ThreatsLynch et al.2025https://arxiv.org/abs/2510.05179
Helpful Assistant Features Suppress Emergent MisalignmentDupré la Tour2025https://alignment.openai.com/helpful-assistant-features/
Sidestepping Evaluation Awareness with Production EvaluationsWilliams et al.2025https://alignment.openai.com/prod-evals/
Teaching Claude WhyKutasov et al.2026https://alignment.anthropic.com/2026/teaching-claude-why/
Positive Alignment: Artificial Intelligence for Human FlourishingLaukkonen et al.2026https://arxiv.org/abs/2605.10310
Concrete Problems in AI SafetyAmodei et al.2016https://arxiv.org/abs/1606.06565
Truthful AI: Developing and governing AI that does not lieEvans et al.2021https://arxiv.org/abs/2110.06674
AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsAndriushchenko et al.2024https://arxiv.org/abs/2410.09024
GPQA: A Graduate-Level Google-Proof Q&A BenchmarkRein et al.2023https://arxiv.org/abs/2311.12022
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Tasks?Deng et al.2025https://arxiv.org/abs/2509.16941