Discussion about this post

User's avatar
fox's avatar
Aug 21Edited

I agree with basically everything here—especially that alignment is largely a systems level problem. However, I’d be cautious about extending principles from human systems directly to AI systems. Our institutions are predicated on a model of agents and agency that doesn’t fully apply to current models, and I’m uncertain whether we’re even on a trajectory toward developing it.

I like to separate the concept into two components: an agent as a functional pattern and an agent as an ontological entity. The functional pattern consists of something interacting with an environment in a loop toward some type of goal whereas the ontological entity is a persistent thing to which goals, interests, actions, and consequences can be attached across time and contexts.

This distinction matters because institutional incentives require a sufficiently persistent entity whose goals can be "brought into alignment". The usual game theoretic implications only apply insofar as they bear on something that has the persistence and unitarity to be responsive to those dynamics. Granted, some of this can be virtualized within functional loops, but in a much thinner and less binding sense than what we typically assume when thinking about human agents.

Where this points, I think, is toward two broad approaches. We can try to make AI agents more persistent, so that they can be governed more readily by human-like institutions and game-theoretic incentives, or we can lean on the more traditional engineering approaches you describe above: validation, checks, monitoring, and so on. I mostly favor the latter, because it seems to fit both what we have now and the direction things are naturally going.

Leslie De Jesus's avatar

The more successful a constraint becomes, the easier it is to forget why we put it there in the first place.

That’s what stayed with me reading this. We can design the harness, constrain the path, add the checks, and still end up optimizing the wrong thing.

I’ve seen the same thing happen in organizations. A process starts by protecting something important. Over time, following the process becomes more important than the reason it exists.

A good test might be to ask less often, “Is this working as designed?” and more often, “Is this still protecting what matters?”

Because perfect compliance can still be misalignment.

7 more comments...

No posts

Ready for more?