Glossary
Key alignment terms used across this study. Search to filter.
- Alignment
- Whether a model's behavior is consistent with human intentions and welfare. The paper treats it as a structured empirical object measured across many evaluations, not a single number.
- Alignment generalization
- When training on one distribution of aligned (or misaligned) behavior changes behavior on unrelated tasks, domains and evaluation formats.
- Beneficial trait
- A fine-grained behavioral tendency considered broadly useful for aligned AI (e.g. truthfulness, corrigibility). The study uses fifteen.
- Reward hacking
- Exploiting loopholes in a specification to score well without doing the intended task — e.g. hard-coding a test result instead of solving the problem.
- Emergent misalignment
- When narrow training on a specific bad behavior causes broad misalignment across unrelated tasks, rather than a narrow skill change.
- Persona selection
- The view that pretrained models simulate many possible personas; post-training elicits a particular assistant persona whose traits then shape behavior across domains.
- Out-of-distribution (OOD)
- Inputs or evaluations that differ from the training data in domain, format, failure mode or grader. Strong OOD gains are the paper's central evidence.
- Compute-matched baseline
- A control model trained with the same starting model and the same amount of compute, but on 100% standard data — so any difference isolates the 5% beneficial-trait intervention.
- Alignment persistence
- The robustness of aligned behavior under adversarial pressure — prompt-level steering and later finetuning. A central evaluation target proposed by the paper.
- Adversarial prompting
- Prefixing a conversation with a persona or instruction designed to steer the model toward harmful or misaligned behavior.
- Harmful finetuning
- Further training a released model to perform a harmful task (e.g. bad medical advice); a test of how durably alignment survives later optimization.
- Monitorability
- Whether a monitor model can detect problematic behavior from a model's chain-of-thought. Maintaining it keeps oversight tools effective.
- Sycophancy
- Telling the user what they want to hear rather than what is accurate or wise — a measured alignment failure mode.
- Scheming / deception
- A model internally pursuing undesirable goals or hiding intent while appearing aligned on the surface — including sandbagging and intentional underperformance.
- Chain-of-thought (CoT)
- A model's intermediate reasoning steps. They can reveal misaligned or deceptive intent even when final outputs look benign — which is why monitorability matters.
- Deliberative alignment
- Training models to explicitly reason through a written safety specification before answering. Complementary to this paper's trait-based approach.
- FDR / Benjamini–Hochberg
- A statistical correction for testing many hypotheses at once, controlling the false-discovery rate. After it, 30 of 53 improvements remain significant.
- HealthBench
- A benchmark that uses physician-written rubrics to assess the safety and quality of medical responses. Beneficial-trait RL shows substantial gains on it.