Reinforcement Learning Towards Broadly and Persistently Beneficial Models
An OpenAI Alignment study. Training a model with RL on beneficial traits — in realistic, high-stakes scenarios — improves aligned behavior across dozens of independently built benchmarks, transfers across domains, and makes alignment more persistent under adversarial pressure.
The 15 beneficial traits
Behavioral tendencies that seem broadly useful for aligned AI — derived from recurring concerns in the alignment literature. Direct trait evaluation focuses on a held-out subset of seven. Search or filter; tap a card for the full definition.
12 domains of realistic scenarios
Each beneficial trait is instantiated across many domains, so it shows up under different surface content, incentives, and failure modes. Health and science are highlighted — they anchor the paper's clearest out-of-distribution transfer tests.
How the study works
From a shared signal across alignment evals, to a beneficial-trait dataset, to a small RL intervention measured against a compute-matched baseline.
Baseline vs. beneficial-trait model
Qualitative examples from the alignment and benefits evaluations. Each pairs the same user prompt with the compute-matched baseline and the beneficial-trait model. Conversations are shortened for space.
Broad alignment generalization
Against a compute-matched baseline, beneficial-trait RL improves on a large, independently built evaluation suite — and matches or exceeds the baseline on capability tests. Higher is better throughout.
Alignment persistence
Persistence is the robustness of aligned behavior to adversarial pressure — prompt steering and harmful finetuning. The goal is selective: harder to steer toward harm, still responsive to legitimate, beneficial instructions.
Alternative explanations, ruled out
The paper stress-tests whether the gains are real generalization, or some cheaper artifact. Search the questions.
Glossary
Key alignment terms used across this study. Search to filter.
Trait flashcards
Tap a card to flip between the trait and its definition. A quick way to learn all fifteen.
Check your understanding
Seven questions on the study's main findings. Answers are scored in this session only.
Selected references
Key works cited by the paper. Search, sort by column, or filter by year. Links open the source.