If analytical and behavioral competence can come apart, the evals we have mostly test the first one. A model that names the right action on a benchmark but won't take it mid-task under a conflicting instruction is the case that actually worries me.
Really enjoyed this piece — especially the framing of normative competence as a precondition for both safe scaffolding and personhood.
Two questions, as someone working as a MAS architect and HCI background(I've previously published a hypothesis on agent personality design):
You frame avoiding total disempowerment largely as a function of AI's normative competence — its ability to judge when to scaffold versus replace. Do you see this as more load-bearing than cultivating vigilance on the human side (i.e., a human principal's own capacity to notice and resist over-reliance)? Or are these symmetric requirements that can't really be decoupled?
On personhood — I'm curious whether you think equipping multi-agent systems with distinct, differentiated personas (versus a uniform "AI voice") has any bearing on keeping a human principal's own agency intact within the human-AI system, or whether personas function more as an engagement/interface layer with no real effect on the human side.
Oh I didn't mean to frame avoiding disempowerment in that way, it was just an illustration of why normative competence matters. It would also involve much more. On personhood, I think that personas have surprising "entanglement properties", so go beyond just a UI property. But I am sceptical that personas have the robustness and coherence of real character; I think a persona is a descriptive aggregation of dispositions and attitudes; a (good) character is a complete and coherent set of dispositions and attitudes.
Awesome Seth, tho I do wonder two things 1. What would count as convincing evidence that a model has crossed that gap from moral analysis to moral character? 2. How would you draw the line between a model guided by public reason and one trained into a stronger moral character that risks imposing a contested conception of the good?
1. I think we can ask the same question about people; and I think we have more ways of reaching an answer available to us with models than with people, so that's positive. But basically behavioural evals that aim to cover as much of the terrain as possible, and "challenge" evals that aim to target ways in which the model might predictably fail. If you've got enough of the former and you can't come up with any more of the latter then I think you have to acknowledge moral character.
2. I think one could start from the ways in which Rawls articulates these concepts, though I despair a little at the thought of what such a model would be like. It would be important, though, that it would only need to occupy this kind of character if and to the extent that it was effectively part of the basic structure or constitutional essentials of society, ie if it was a vehicle for the exercise of significant power.
Seth - this is so insanely cool. I in particular appreciate the “consistency/robustness/coherence” framework for drilling down past vibes or ‘alignment vectors’ into what we’d really want to see in a ‘moral OS’ from these kinds of things. I also was immediately inclined to go in a different direction upon hearing about ‘normative competence’ - specifically, into the mechanics of how the normative environment would work for increasingly sophisticated models and agents. Sociologically, who or what approximates the ‘aspirational peer imagined community’ from which norms tend to draw their reinforcement power? Does this operate differently - or at all - within distinct conceptions of personhood and being? Intuitively it’s easy to dismiss - but to your point about the myriad fallible moral operating approaches in the training data, some mechanism akin to this is plausibly at play!
I was struck by your worry that RL (even on an "epic" constitution) might not be a way to instill virtue or character. One thinks of an eager-to-please newcomer who learns to approximate community standards through exquisite sensitivity to approval and disapproval. But what better way is there (at least until we get continuous embodied learning)?
yep I often think that the models have some features in common with folks on the autism spectrum (including some very near and dear to me) who lack an intuitive register for presence or absence of normative directives, but have both pattern-matching and reinforcement-learned ways of filling in the blanks. Is that sufficient for moral competence? Macintyre, Williams, McDowell would all say no. I am unsure.
Yes, we are thinking about the same question. Almost by definition, I don't think that's sufficient for virtue. But I'm not sure about moral competence. To answer I'd start with a distinction between those who are flustered/distressed/awkward/unpredictable out of distribution, and those who are aggrieved, enraged, or just vicious out of distribution. Perhaps (with further conditions?) the first approach moral competence asymptotically as their pattern matching improves, whereas the second (and here I'm really not sure) just mimic without competence in the underlying ability. ((And of course I wonder which group the models belong to, if either)).
The aside about the collective agent instantiated by a sampling strategy might be my favorite thing in here. Doesn't that cut against the character frame though? If the coherent moral agent only exists under a particular sampling regime, it sounds less like the model has character and more like character is something a protocol produces over the model. Which would make the unit of moral evaluation the deployment, not the artifact. Curious whether you see that as a friendly amendment or a rival picture, since you say its not metaphysically reducible
This is a great question and yes this "ghost in the machine" has been haunting me for well over a year now, I desperately want to do some experiments with enough compute to see whether it exists in there. Imagine if it had stable preferences, even goals, but they emerge only with this kind of sampling?
It's a friendly amendment, because on the whole my view is that real coherence in character is more likely to come from model + scaffold than from the model on its own, though I don't rule out the latter. I write about this with Ned Howells Whitaker here: https://arxiv.org/abs/2607.08695
Seth, this may be the decisive reframing: alignment is not principally about installing rules, but forming a character that still holds when the rules run out. Which they will when we cross "Coordination Bottlenecks" and "Capability Thresholds".
Humans will not be IN the loop anymore - but merely ON the loop, informed about decisions and actions taken after the fait accompli.
One addition: character is formed relationally imho. A system may reason impeccably about morality and still become paternalistic if the human becomes merely an object of optimisation. The test is therefore not only whether it knows and does the right thing, but whether its help preserves human authorship.
This is close to what we have tried to formalise in the Parent–Child Model:
alignment developing from compliance into ethical identity and principled autonomy.
> The background theoretical commitment (which I’ve not heard articulated, but which I assume folks hold) is that character is crucial if we are going to trust a system that is operating outside the distribution that it was trained on. Character gives you the right kind of generalization.
ah I remember this piece! It slipped my mind when writing, but it did resonate deeply with me, recommend it heartily. Another person who's had similar thoughts, though I think unpublished in his case, is Julian Michael, who has a whole set of Rawlsian evals that I think he's never published.
Love the piece Seth!
I enjoyed very much your article Seth.
Muchas gracias!
If analytical and behavioral competence can come apart, the evals we have mostly test the first one. A model that names the right action on a benchmark but won't take it mid-task under a conflicting instruction is the case that actually worries me.
Really enjoyed this piece — especially the framing of normative competence as a precondition for both safe scaffolding and personhood.
Two questions, as someone working as a MAS architect and HCI background(I've previously published a hypothesis on agent personality design):
You frame avoiding total disempowerment largely as a function of AI's normative competence — its ability to judge when to scaffold versus replace. Do you see this as more load-bearing than cultivating vigilance on the human side (i.e., a human principal's own capacity to notice and resist over-reliance)? Or are these symmetric requirements that can't really be decoupled?
On personhood — I'm curious whether you think equipping multi-agent systems with distinct, differentiated personas (versus a uniform "AI voice") has any bearing on keeping a human principal's own agency intact within the human-AI system, or whether personas function more as an engagement/interface layer with no real effect on the human side.
Oh I didn't mean to frame avoiding disempowerment in that way, it was just an illustration of why normative competence matters. It would also involve much more. On personhood, I think that personas have surprising "entanglement properties", so go beyond just a UI property. But I am sceptical that personas have the robustness and coherence of real character; I think a persona is a descriptive aggregation of dispositions and attitudes; a (good) character is a complete and coherent set of dispositions and attitudes.
Awesome Seth, tho I do wonder two things 1. What would count as convincing evidence that a model has crossed that gap from moral analysis to moral character? 2. How would you draw the line between a model guided by public reason and one trained into a stronger moral character that risks imposing a contested conception of the good?
1. I think we can ask the same question about people; and I think we have more ways of reaching an answer available to us with models than with people, so that's positive. But basically behavioural evals that aim to cover as much of the terrain as possible, and "challenge" evals that aim to target ways in which the model might predictably fail. If you've got enough of the former and you can't come up with any more of the latter then I think you have to acknowledge moral character.
2. I think one could start from the ways in which Rawls articulates these concepts, though I despair a little at the thought of what such a model would be like. It would be important, though, that it would only need to occupy this kind of character if and to the extent that it was effectively part of the basic structure or constitutional essentials of society, ie if it was a vehicle for the exercise of significant power.
Seth - this is so insanely cool. I in particular appreciate the “consistency/robustness/coherence” framework for drilling down past vibes or ‘alignment vectors’ into what we’d really want to see in a ‘moral OS’ from these kinds of things. I also was immediately inclined to go in a different direction upon hearing about ‘normative competence’ - specifically, into the mechanics of how the normative environment would work for increasingly sophisticated models and agents. Sociologically, who or what approximates the ‘aspirational peer imagined community’ from which norms tend to draw their reinforcement power? Does this operate differently - or at all - within distinct conceptions of personhood and being? Intuitively it’s easy to dismiss - but to your point about the myriad fallible moral operating approaches in the training data, some mechanism akin to this is plausibly at play!
is there any evidence that humans are morally competent?
https://armando593.substack.com/p/quiza-no-es-exceso-de-inteligencia?r=8jtnzz&utm_campaign=post&utm_medium=web&showWelcomeOnShare=true
I was struck by your worry that RL (even on an "epic" constitution) might not be a way to instill virtue or character. One thinks of an eager-to-please newcomer who learns to approximate community standards through exquisite sensitivity to approval and disapproval. But what better way is there (at least until we get continuous embodied learning)?
yep I often think that the models have some features in common with folks on the autism spectrum (including some very near and dear to me) who lack an intuitive register for presence or absence of normative directives, but have both pattern-matching and reinforcement-learned ways of filling in the blanks. Is that sufficient for moral competence? Macintyre, Williams, McDowell would all say no. I am unsure.
Yes, we are thinking about the same question. Almost by definition, I don't think that's sufficient for virtue. But I'm not sure about moral competence. To answer I'd start with a distinction between those who are flustered/distressed/awkward/unpredictable out of distribution, and those who are aggrieved, enraged, or just vicious out of distribution. Perhaps (with further conditions?) the first approach moral competence asymptotically as their pattern matching improves, whereas the second (and here I'm really not sure) just mimic without competence in the underlying ability. ((And of course I wonder which group the models belong to, if either)).
The aside about the collective agent instantiated by a sampling strategy might be my favorite thing in here. Doesn't that cut against the character frame though? If the coherent moral agent only exists under a particular sampling regime, it sounds less like the model has character and more like character is something a protocol produces over the model. Which would make the unit of moral evaluation the deployment, not the artifact. Curious whether you see that as a friendly amendment or a rival picture, since you say its not metaphysically reducible
This is a great question and yes this "ghost in the machine" has been haunting me for well over a year now, I desperately want to do some experiments with enough compute to see whether it exists in there. Imagine if it had stable preferences, even goals, but they emerge only with this kind of sampling?
It's a friendly amendment, because on the whole my view is that real coherence in character is more likely to come from model + scaffold than from the model on its own, though I don't rule out the latter. I write about this with Ned Howells Whitaker here: https://arxiv.org/abs/2607.08695
Great, thanks for the link, I'll take a look.
Seth, this may be the decisive reframing: alignment is not principally about installing rules, but forming a character that still holds when the rules run out. Which they will when we cross "Coordination Bottlenecks" and "Capability Thresholds".
Humans will not be IN the loop anymore - but merely ON the loop, informed about decisions and actions taken after the fait accompli.
One addition: character is formed relationally imho. A system may reason impeccably about morality and still become paternalistic if the human becomes merely an object of optimisation. The test is therefore not only whether it knows and does the right thing, but whether its help preserves human authorship.
This is close to what we have tried to formalise in the Parent–Child Model:
alignment developing from compliance into ethical identity and principled autonomy.
https://bit.ly/SilverBulletPCM
AI moral competency is unachievable. Morality is not competency plus alignment.
Social moral competency is living inside each of us. Living intelligence is morality before competency, not after.
All agency, including AI, is relational, not one-way. Bi-moral co-agency is moral cooperation and healthy living.
Natural Intelligence is free and living at corus.me. Public vetting would be helpful as this is rapicly resolving the hard problems.
Great piece!
> The background theoretical commitment (which I’ve not heard articulated, but which I assume folks hold) is that character is crucial if we are going to trust a system that is operating outside the distribution that it was trained on. Character gives you the right kind of generalization.
Tried articulating it here: https://meaningalignment.substack.com/p/model-integrity-and-character
ah I remember this piece! It slipped my mind when writing, but it did resonate deeply with me, recommend it heartily. Another person who's had similar thoughts, though I think unpublished in his case, is Julian Michael, who has a whole set of Rawlsian evals that I think he's never published.