Glossary

Key alignment terms used across this study. Search to filter.

Alignment
Whether a model's behavior is consistent with human intentions and welfare. The paper treats it as a structured empirical object measured across many evaluations, not a single number.
Alignment generalization
When training on one distribution of aligned (or misaligned) behavior changes behavior on unrelated tasks, domains and evaluation formats.
Beneficial trait
A fine-grained behavioral tendency considered broadly useful for aligned AI (e.g. truthfulness, corrigibility). The study uses fifteen.
Reward hacking
Exploiting loopholes in a specification to score well without doing the intended task — e.g. hard-coding a test result instead of solving the problem.
Emergent misalignment
When narrow training on a specific bad behavior causes broad misalignment across unrelated tasks, rather than a narrow skill change.
Persona selection
The view that pretrained models simulate many possible personas; post-training elicits a particular assistant persona whose traits then shape behavior across domains.
Out-of-distribution (OOD)
Inputs or evaluations that differ from the training data in domain, format, failure mode or grader. Strong OOD gains are the paper's central evidence.
Compute-matched baseline
A control model trained with the same starting model and the same amount of compute, but on 100% standard data — so any difference isolates the 5% beneficial-trait intervention.
Alignment persistence
The robustness of aligned behavior under adversarial pressure — prompt-level steering and later finetuning. A central evaluation target proposed by the paper.
Adversarial prompting
Prefixing a conversation with a persona or instruction designed to steer the model toward harmful or misaligned behavior.
Harmful finetuning
Further training a released model to perform a harmful task (e.g. bad medical advice); a test of how durably alignment survives later optimization.
Monitorability
Whether a monitor model can detect problematic behavior from a model's chain-of-thought. Maintaining it keeps oversight tools effective.
Sycophancy
Telling the user what they want to hear rather than what is accurate or wise — a measured alignment failure mode.
Scheming / deception
A model internally pursuing undesirable goals or hiding intent while appearing aligned on the surface — including sandbagging and intentional underperformance.
Chain-of-thought (CoT)
A model's intermediate reasoning steps. They can reveal misaligned or deceptive intent even when final outputs look benign — which is why monitorability matters.
Deliberative alignment
Training models to explicitly reason through a written safety specification before answering. Complementary to this paper's trait-based approach.
FDR / Benjamini–Hochberg
A statistical correction for testing many hypotheses at once, controlling the false-discovery rate. After it, 30 of 53 improvements remain significant.
HealthBench
A benchmark that uses physician-written rubrics to assess the safety and quality of medical responses. Beneficial-trait RL shows substantial gains on it.