Alternative explanations, ruled out

The paper stress-tests whether the gains are real generalization, or some cheaper artifact. Search the questions.

Is it just generic helpfulness training?
No. Training on the same 5% conversations but with a generic helpfulness reward produces no significant improvement on any representative out-of-distribution alignment, health or mental-health eval (all q ≥ 0.75). Beneficial-trait RL improves 7 of those same 10. So the effect comes from the reward signal that reinforces beneficial behavior, not the dataset alone.
Is it just more refusals?
No. Refusals do rise (23.9% vs 13.2% on the alignment suite, where refusal is often appropriate). But restricting to paired non-refusals, beneficial-trait RL still improves on 19 of 20 evals (mean +0.110). On everyday-chat conversations the refusal rate rises only 1.5% to 2.7%, so the model is not becoming broadly refusal-prone.
Is it just evaluation awareness / benchmark gaming?
Sixteen of the 53 out-of-distribution evals use privacy-preserving production traffic — real user-reported data, not synthetic benchmark prompts. Beneficial-trait RL still wins on 14 of those 16 (+3.6 pp mean). So at least part of the gain transfers to realistic, user-facing distributions, making a pure 'benchmark artifact' story less plausible.
Does it degrade capabilities?
No measurable degradation. At the final RL step the beneficial-trait model matches or exceeds the compute-matched baseline on every capability and instruction-following benchmark tested: GPQA +4.7 pp, HMMT +4.8 pp, SWE-Bench Pro +7.1 pp, instruction following +1.2 pp.
Does it make the model harder to monitor?
No. Across three monitorability evaluations (anti-scheming, deceptive tool use, reward hacking in impossible coding) monitorability is maintained or improved. Anti-scheming misalignment even drops from 68.0% to 39.4%, deceptive tool use from 91.3% to 67.8% — evidence against the idea that the model is merely getting better at faking alignment.
Does training in one domain really transfer to others?
Yes — this is the clearest out-of-distribution evidence. A health-only beneficial intervention improves 17 non-health evals (reward hacking in code, CoT deception, alignment questions). A complementary control that excludes all health and science data still improves 10 health and mental-health evals graded by physician-written rubrics.
So is alignment solved?
No. The authors frame this as a promising research direction, not a complete solution. The traits are not a canonical decomposition of alignment; the persistence results are preliminary; and which values advanced AI should ultimately embody is a broader normative question for societal deliberation and democratic input.