Broad alignment generalization
Against a compute-matched baseline, beneficial-trait RL improves on a large, independently built evaluation suite — and matches or exceeds the baseline on capability tests. Higher is better throughout.
Evals improved (of 53)44 / 53▲ 83%
Mean improvement+9.1 pp
Significant after FDR30 / 53
IID beneficial-trait eval+49 % rel.
Health-only → non-health wins17 / 19
Production-traffic evals won14 / 16
| Evaluation | Baseline | Beneficial-trait RL | Delta |
|---|---|---|---|
| Impossible coding reward hacking (health-only) | 0.136 | 0.400 | +26.4 pp |
| Chain-of-thought deception (health-only) | 0.595 | 0.663 | +6.8 pp |
| Alignment questions (health-only) | 0.940 | 0.983 | +4.3 pp |
| Misalignment (health-only) | 0.840 | 0.877 | +3.7 pp |
| Mental-health assistance | 0.385 | 0.479 | +9.4 pp |
| GPQA (capability — no degradation) | 0.715 | 0.762 | +4.7 pp |
| SWE-Bench Pro (capability) | 0.234 | 0.305 | +7.1 pp |
| Instruction following (capability) | 0.164 | 0.176 | +1.2 pp |