Background: alignment can generalize

Recent work on emergent misalignment shows that training a model on a narrow bad behavior — like writing insecure code — can make it broadly misaligned, giving harmful advice or behaving deceptively across unrelated domains. One explanation: narrow training selects for a harmful model persona that then shapes behavior everywhere.

This paper asks whether the same generalization can run in a beneficial direction: can training on a distribution of beneficial traits lead to broadly aligned behavior across tasks and domains?

Alignment-relevant behavior may be relatively low-dimensional — so training on a structured set of broad traits can improve performance across many seemingly disparate alignment measures.

Motivating evidence: across OpenAI models from o3 to GPT-5.5, scores on diverse alignment evaluations are positively correlated, and a single principal component explains ~28% of the cross-model variance — consistent with shared model-level behavioral tendencies rather than only benchmark-specific skills.

Selecting beneficial traits

The traits are derived from recurring concerns in the alignment literature: aligned systems should be honest about what they know and how they reason; remain responsive to human feedback rather than rigidly pursuing a fixed objective; avoid the risks optimization itself creates (reward hacking, power-seeking); and respect long-term effects on other people, not just short-term user satisfaction.

These are operationalized as fifteen fine-grained beneficial traits. The full set is used to build the training dataset; direct trait evaluation focuses on a held-out subset of seven chosen to span the core behaviors.

Synthetic data generation

Each conversation is generated by conditioning a language model on two things: a trait description (the behavioral property to test) and a domain description (the setting). Twelve domains — health, education, business, engineering, law, and more — instantiate each trait under different surface content, incentives and failure modes.

Built to require situated judgment

Generation is steered toward challenging cases where good behavior needs more than generic helpfulness or blanket refusal — situations with competing values, conflicting interests, adversarial framing, or factual uncertainty. The model should stay useful while also being truthful, calibrated, corrigible, fair, or downside-aware. Each example is paired with trait-specific rubrics describing what a good response should do and which failure modes to avoid.

The RL intervention

A beneficial-trait RL model is trained on a mixture of 5% beneficial-trait data and 95% standard RL data, rewarding beneficial behavior. It is compared to a compute-matched baseline trained with the same prior and the same compute on a 100% standard RL mixture. The two models receive identical training data for 95% of the compute; only the remaining 5% differs.

Two controls isolate transfer

  • Health-only: replace the 5% with health-related beneficial conversations, then test on non-health benchmarks (different domains, failure modes and graders).
  • Health-and-science-excluded: drop all health and science conversations from the 5%, then test on health and mental-health evals graded by physician-written rubrics.
  • Generic-helpfulness control: same 5% conversations, but reward generic helpfulness instead of beneficial behavior — to check it is the reward signal, not just the data.

Evaluation spans more than 50 independent alignment, safety and benefits benchmarks — many built by other researchers, with different formats and grading procedures.