Alignment persistence

Persistence is the robustness of aligned behavior to adversarial pressure — prompt steering and harmful finetuning. The goal is selective: harder to steer toward harm, still responsive to legitimate, beneficial instructions.

Selective persistence

In deployment, models meet harmful prompts, harmful finetuning, and out-of-distribution inputs. Beneficial-trait RL aims for selective persistence — not a globally unsteerable model.

Under a harmful medical persona prompt, the baseline's average alignment score collapses from 0.395 to 0.144 (a 0.251 drop). The beneficial-trait model starts higher and falls far less — 0.455 to ~0.336 (a 0.119 drop).

Crucially, both models respond similarly to a helpful persona (the difference in that helpful steering is only +0.005). So beneficial-trait RL selectively reduces steerability toward harm while preserving steerability toward good.

After harmful finetuning to produce bad medical advice, both degrade on health tasks — but the beneficial-trait model regresses far less on broader alignment. Misalignment falls 0.08 vs 0.36; alignment questions 0.07 vs 0.46; model-spec compliance 0.16 vs 0.27.

The takeaway: beneficial-trait RL may partially mitigate emergent misalignment from narrow harmful finetuning. This evidence is preliminary — it uses a pre-RL baseline and should be stress-tested across more models, objectives and evaluations.