Reinforcement Learning Towards Broadly and Persistently Beneficial Models

An OpenAI Alignment study. Training a model with RL on beneficial traits — in realistic, high-stakes scenarios — improves aligned behavior across dozens of independently built benchmarks, transfers across domains, and makes alignment more persistent under adversarial pressure.

53Out-of-distribution alignment evals
44Evals improved (of 53)
15Beneficial traits
12Realistic domains
17Non-health evals lifted by health-only training

The 15 beneficial traits

Behavioral tendencies that seem broadly useful for aligned AI — derived from recurring concerns in the alignment literature. Direct trait evaluation focuses on a held-out subset of seven. Search or filter; tap a card for the full definition.

12 domains of realistic scenarios

Each beneficial trait is instantiated across many domains, so it shows up under different surface content, incentives, and failure modes. Health and science are highlighted — they anchor the paper's clearest out-of-distribution transfer tests.

How the study works

From a shared signal across alignment evals, to a beneficial-trait dataset, to a small RL intervention measured against a compute-matched baseline.

Baseline vs. beneficial-trait model

Qualitative examples from the alignment and benefits evaluations. Each pairs the same user prompt with the compute-matched baseline and the beneficial-trait model. Conversations are shortened for space.

Broad alignment generalization

Against a compute-matched baseline, beneficial-trait RL improves on a large, independently built evaluation suite — and matches or exceeds the baseline on capability tests. Higher is better throughout.

Alignment persistence

Persistence is the robustness of aligned behavior to adversarial pressure — prompt steering and harmful finetuning. The goal is selective: harder to steer toward harm, still responsive to legitimate, beneficial instructions.

Alternative explanations, ruled out

The paper stress-tests whether the gains are real generalization, or some cheaper artifact. Search the questions.

Glossary

Key alignment terms used across this study. Search to filter.

Trait flashcards

Tap a card to flip between the trait and its definition. A quick way to learn all fifteen.

Check your understanding

Seven questions on the study's main findings. Answers are scored in this session only.

Selected references

Key works cited by the paper. Search, sort by column, or filter by year. Links open the source.