Check your understanding

Seven questions on the study's main findings. Answers are scored in this session only.

Score 0 / 7

1What is the core claim of the paper?

2On how many of the 53 OOD alignment evals did beneficial-trait RL beat the compute-matched baseline?

3What does the health-only experiment demonstrate?

4What does 'selective persistence' mean here?

5Why train a 'generic-helpfulness' control on the same data?

6Did beneficial-trait RL hurt capabilities or monitorability?

7How do the authors frame the result overall?