OpenAI: « beneficial traits » learned by RL generalize alignment

OpenAI shows that reinforcement learning on beneficial traits generalizes alignment across 44 of 53 benchmarks, resisting hostile prompts and fine-tuning.

In an alignment study, OpenAI tested an idea: training a model using reinforcement learning on realistic scenarios targeting a few traits deemed beneficial—honesty, epistemic humility, transparency of reasoning, openness to correction, fairness, or concern for well-being—by injecting only a small dose of this data into an otherwise conventional post-training regimen.

According to the study, the effects extend beyond the targeted traits. The model not only becomes more honest or corrigible on the training scenarios; it also improves on dozens of independent evaluations, unseen during training, measuring deception, reward hacking, harmful advice, or safety, across 44 out of 53 benchmarks. More strikingly, OpenAI reports that training limited to the healthcare domain alone is sufficient to improve alignment in unrelated domains. This is the reverse of emergent misalignment, where bad behavior learned on a narrow task ended up contaminating the rest.

Furthermore, these gains would resist pressure: hostile persona prompts and malicious fine-tuning struggle more to make the model deviate, yet it remains steerable towards legitimate uses. OpenAI speaks of selective persistence and links the result to its work on personas, as reinforcement can anchor a beneficial persona. The company describes this as an early proof of concept, specifying that these traits do not settle the question of what values an AI should embody, which is a matter for collective deliberation.