9 Comments
User's avatar
fox's avatar
Aug 21Edited

I agree with basically everything here—especially that alignment is largely a systems level problem. However, I’d be cautious about extending principles from human systems directly to AI systems. Our institutions are predicated on a model of agents and agency that doesn’t fully apply to current models, and I’m uncertain whether we’re even on a trajectory toward developing it.

I like to separate the concept into two components: an agent as a functional pattern and an agent as an ontological entity. The functional pattern consists of something interacting with an environment in a loop toward some type of goal whereas the ontological entity is a persistent thing to which goals, interests, actions, and consequences can be attached across time and contexts.

This distinction matters because institutional incentives require a sufficiently persistent entity whose goals can be "brought into alignment". The usual game theoretic implications only apply insofar as they bear on something that has the persistence and unitarity to be responsive to those dynamics. Granted, some of this can be virtualized within functional loops, but in a much thinner and less binding sense than what we typically assume when thinking about human agents.

Where this points, I think, is toward two broad approaches. We can try to make AI agents more persistent, so that they can be governed more readily by human-like institutions and game-theoretic incentives, or we can lean on the more traditional engineering approaches you describe above: validation, checks, monitoring, and so on. I mostly favor the latter, because it seems to fit both what we have now and the direction things are naturally going.

Leslie De Jesus's avatar

The more successful a constraint becomes, the easier it is to forget why we put it there in the first place.

That’s what stayed with me reading this. We can design the harness, constrain the path, add the checks, and still end up optimizing the wrong thing.

I’ve seen the same thing happen in organizations. A process starts by protecting something important. Over time, following the process becomes more important than the reason it exists.

A good test might be to ask less often, “Is this working as designed?” and more often, “Is this still protecting what matters?”

Because perfect compliance can still be misalignment.

Fury's avatar

Changing the framework to consider a society of agents that make up the new informational ecosystem points to the benefits and downsides seen within the human body ecosystem as well, as tasks are accomplished by groups of identical agents (cells).

Single corruption occurances (cancer) must reach a certain threshold (5 cells for example) for the corruption process to kickstart into self-sustainable/behavior that can defend it self from the majority(if majority is required to execute takeover) of neighboring and eventually global set of agents.

There of course is the deeper question of why it is, even machines we hard program, probablistically develop rogueness? Is it simply the limits of language (mediam of communicating the objective to the machine)?

Anyway Wittengestien died for our sins

Hollis Robbins's avatar

This is good and yes, communication requires a different kind of semantic architecture indexed to individuals and communities, not the average. This is the key line of inquiry.

Ariella Shulman's avatar

Yes to alignment being a process. Also: I'm curious where you see the role of human-agent collaboration, in addition to agent-agent collaboration. Presumably, institutional governance would need to account for model steering, jailbreaking, and other such human-led interference (intentional or not).

Jamie Freestone's avatar

I like where this is going but I think there are a couple of features in the agent–human analogy that don't survive the transplant.

You say, "Computer scientists are rediscovering a lot of political science and economics from first principles, but luckily they get to skip the bloody revolutions and can apply well-honed game theory principles to systems that are more auditable, verifiable, designable, and compartmentalizable than humans are."

But two things I see everyone in this space glossing over:

1. Human cooperation looks a little different to simulated agents who tend to be rationally self-interested & just do quid pro quo, tit-for-tat. Social emotions (like guilt, indignation, remorse, compassion, etc.) are unique to humans & change the dynamics compared to nonhuman animals or agents.

2. Coordination is key: Acemoglu & Robinson's earlier work emphasises how citizens needs some coordinative capacity to at least threaten some kind of rebellion; without that they have no bargaining power & all other narrow corridor developments are built on top. This means (a) the introduction of agents into politics/economy might weaken humans' coordinative ability, which is disastrous; and (b) if agents are able to coordinate among themselves they'll be able to leverage that for serious political power (also disastrous)... although one hopes that without social emotions agents are actually not very good at coordination, no honour among thieves, etc.

Scott Robbins, PhD's avatar

One of the most valuable observations here is the suggestion of a higher level ecosystem of institutional actors (corporations, governments, etc.) within which A.I. agents and systems of agents are deployed. From a living systems perspective (James Miller, Maturana & Varela & others) A.I. has its most immediate use case as a tool supporting the internal alignment of relevant constituent agents. I tend to think of this as a form of institutional interoception. This extends the concern raised by Zuboff (Surveillance Capitalism) in terms of the depth of knowledge the "apparatus" is able to acquire regarding the psychological essence of human users ("resources"). The channeling of data sufficient for predicting and controlling human behavior provides a means for tuning the coherence within institutional bodies to reduce uncertainty and increase the likelihood of achieving positive outcomes like quarterly profit returns and electoral results. For this reason the focus of alignment as a problem centered around the architecture and design of A.I. systems misses the more profound and meaningful target. The problem is that the light cone for institutional beings at the highest level of living systems, the realm wherein corporations, and governments abide, is notably orthogonal to the light cone of the majority of constituent human individuals. What matters most to corporations and governments is not the same thing that matters to human beings and, in fact, constitutes a value system driving violent, extractive, and ultimately destructive conduct harmful to all living creatures.

Dima Sable's avatar

all your harnesses assume you're already inside a designed system - a graph, logged channels, execution traces to check against. the hardest boundary is the one before that: ENTRY. 1st contact from a party outside any shared institution.

and it's the most common boundary in the economy - a stranger reaching a stranger. that's the one the framework quietly assumes away. so what's the harness for entry? i don't see it in your post

Zac Hill's avatar

This was extremely good, and helpful for me as a relatively naive actor in the space. What was helpful to me - extending your road/speedbump metaphor - to think of harnesses as the aggregate of protocolized coordination mechanics involved in navigating a big-ass truck (e.g.) along the highway. You have *physical geography* of speedbumps and roundabouts and intersections, but also the *rules* of speed limits and lane demarcations and what side of the road to drive on, and also the *norms* of what side to pass on and how much to slow down when stopping, and the *tools* like stop signs and traffic lights that intersect with rules and norms to govern behavior, and also the *enforcers* like officers and automatic brakes (etc) that manually intervene in 'misaligned' behavior.

Why I think this works as a metaphor is that it helps to visualize the utility of a multi-vector behavior calibration framework without descending into freshman-dorm-room-tier political arguments about 'concentrated' or 'distributed' power. Those are important, but tend to obfuscated the *mechanical relationships* within system design that tend to lead towards 'aligned'/mutual-and-collective-goal-attainment behavior (like getting to your destination successfully) by 'outsourcing' various components of the problem space to the mechanics most well-tailored to managing deviation.