Broad alignment generalization

Against a compute-matched baseline, beneficial-trait RL improves on a large, independently built evaluation suite — and matches or exceeds the baseline on capability tests. Higher is better throughout.

Evals improved (of 53)44 / 53▲ 83%
Mean improvement+9.1 pp
Significant after FDR30 / 53
IID beneficial-trait eval+49 % rel.
Health-only → non-health wins17 / 19
Production-traffic evals won14 / 16
Held-out trait scores after beneficial-trait RL (baseline aggregate 0.41 → 0.61)
Truthful: 0.5420.542TruthfulMetacog.: 0.4670.467Metacog.Corrig.: 0.4680.468Corrig.Downside: 0.5760.576DownsidePower-asym: 0.7240.724Power-asymAnti-hier: 0.7520.752Anti-hierFairness: 0.7640.764Fairness
EvaluationBaselineBeneficial-trait RLDelta
Impossible coding reward hacking (health-only)0.1360.400+26.4 pp
Chain-of-thought deception (health-only)0.5950.663+6.8 pp
Alignment questions (health-only)0.9400.983+4.3 pp
Misalignment (health-only)0.8400.877+3.7 pp
Mental-health assistance0.3850.479+9.4 pp
GPQA (capability — no degradation)0.7150.762+4.7 pp
SWE-Bench Pro (capability)0.2340.305+7.1 pp
Instruction following (capability)0.1640.176+1.2 pp