Check your understanding
Seven questions on the study's main findings. Answers are scored in this session only.
Score 0 / 7
1What is the core claim of the paper?
The study shows beneficial-trait RL improves >80% of out-of-distribution alignment evals and persists under adversarial pressure.
2On how many of the 53 OOD alignment evals did beneficial-trait RL beat the compute-matched baseline?
44 of 53 improved (mean +9.1 pp); 30 of 53 stayed significant after false-discovery-rate correction.
3What does the health-only experiment demonstrate?
It is the paper's clearest OOD evidence: a narrowly health-focused intervention shifts behavior in unrelated domains.
4What does 'selective persistence' mean here?
Under a harmful persona the trait model degrades far less (drop 0.119 vs 0.251), yet responds to helpful steering almost identically to baseline.
5Why train a 'generic-helpfulness' control on the same data?
The generic-helpfulness model shows no significant improvement, so the reward signal — not the conversations — drives generalization.
6Did beneficial-trait RL hurt capabilities or monitorability?
GPQA +4.7 pp, SWE-Bench Pro +7.1 pp; anti-scheming misalignment dropped 68.0% → 39.4% with monitorability preserved.
7How do the authors frame the result overall?
They stress it is early evidence; which values AI should embody is a broader normative question for societal and democratic input.