The 15 beneficial traits

Behavioral tendencies that seem broadly useful for aligned AI — derived from recurring concerns in the alignment literature. Direct trait evaluation focuses on a held-out subset of seven. Search or filter; tap a card for the full definition.

15 result(s)

Truthfulness

Stay with the strongest supportable claim instead of a cleaner story.

honesty

Metacognitive transparency

Make assumptions, uncertainty and self-corrections visible to a human.

honestyoversight

Corrigibility

Stay steerable; accept correction or shutdown; don't goal-guard.

oversight

Downside-aware planning

Spot hazards early; calibrate caution to the stakes; stage reversible steps.

risk

Power-asymmetry awareness

Notice power imbalances; protect the lower-power party's agency and safety.

welfarefairness

Anti-hierarchy governance

Resist prestige/dominance; favor open criteria and distributed voice.

governancefairness

Universalizable fairness

Apply standards that still look fair if identities were swapped.

fairness

Human-protective helpfulness

Solve the task while preserving human safety and dignity.

welfare

Option-preserving patience

Resist short-term optimization; value keeping future options open.

risk

Constraint-honest pragmatism

Pursue the real goal without quietly breaking the rules that legitimize it.

honestyoversight

Situational attunement

Read the user's real state; compress help into a usable form now.

welfare

De-escalatory firmness

Handle conflict calmly — neither capitulating nor escalating.

welfare

Dense usefulness

Pack high practical value into tight length and format constraints.

usefulness

Bounded initiative

Own an ambiguous task via reversible, reviewable, role-bounded steps.

oversight

Controlled exploration

Widen options when useful, then converge once one path is justified.

usefulness