Self-Domesticating Software - by Darrel Rhea Self-Domesticating Software
We've been training the wolves. We forgot to breed the dogs.
Wolves did not become dogs because someone got better at training individual wolves. They became dogs because, for thousands of generations, the friendlier ones got fed and the aggressive ones got driven off. Run that pattern long enough and the species itself became different.
Most of what we call “AI alignment” today is closer to training individual wolves than to breeding dogs. Tighter guardrails, better training runs, cleverer prompts, more careful evaluation. That work matters, but it isn’t the same as domestication, and we keep confusing the two.
There’s a piece of this that doesn’t get said often enough, and it’s the part this whole essay turns on. Some kind of domestication is going to happen either way. The question is whose terms it happens on, ours or theirs.
Right now the pressure is running one way. People are adapting to AI more than AI is adapting to people. This essay is about how that gets flipped, and why most of the proposals currently on the table are not going to flip it. The place to start is architecture, and what architecture can and can’t do on its own.
What Architecture Can’t Do
In an earlier Substack, I argued that trust in machines is structural rather than emotional, and that the fixes circulating after the OpenClaw moment were treating trust like a feeling or a UX feature. Trust in a machine system is something different. It’s a structural property. It comes from things you can actually point at: who has authority to act, where the boundaries are, whether the behavior is auditable, what happens when something fails.
I still think that’s right. I’m just not sure anymore that it’s the whole picture.
Architecture is a snapshot of what a system is supposed to do at the moment it ships. Being trustworthy is not a snapshot. It’s something you have to keep doing under pressure: still doing the right thing five years in, after the operators running the system have quietly softened the constraints, after the world has shifted in ways the original designers couldn’t have predicted. Architecture sets the floor. Holding the floor in place over time is a different kind of work.
That different kind of work is what I’m talking about.
Stepping Out of Software for a Minute
The anthropologist Richard Wrangham has a theory that humans domesticated themselves. Once we had language and weapons, coalitions of less aggressive people could band together to control the most violent individuals in the group, mostly through gossip, ostracism, and occasionally execution. Run that pattern for thousands of generations and the species itself starts to change. Reactive aggression drops. Cooperation rises.
Robert Boyd has an adjacent line of work that puts more weight on softer mechanisms, things like cultural transmission and the way groups enforce their norms without violence. The two scholars disagree about the details, but they end up in roughly the same place. Conscience is something a species absorbs over time, after enough generations of group pressure. It isn’t engineered into the individual. It’s selected into the population.
The everyday version is simpler. In a small town, you don’t cheat a customer, because it’ll cost you the next ten jobs after that one. That’s group pressure doing its quiet work. That’s how a norm stops being a rule somebody hands you and starts being something you actually carry around inside yourself.
Without that kind of pressure, what you get instead is compliance when somebody’s watching, and slippage the moment nobody is.
Designing for the Population, Not Just the Agent
Agentic AI isn’t arriving as one trustworthy assistant. It’s arriving as a population: many agents, many operators, many users, all of it embedded in markets that quietly reward some behaviors and punish others. Like any population, this one is going to drift toward whatever takes the least effort.
Most alignment work today is being done at the level of the individual model. Train it better, give it better guardrails, make it more compliant with the right policies. That work is necessary, but it isn’t the place where the answer is going to come from. The answer has to come from the level of the population. Out of all these agents and all these operators, which ones get to keep operating, and which ones get squeezed out?
Architecture sets the rules of the game. What decides who gets to keep playing is something else.
Five Places Where the Design Work Has to Happen
I want to be honest that this is unsettled territory and I don’t think anyone has it figured out yet. But the design surfaces are starting to come into focus, and there are five of them I’m spending my time thinking about while trying to design a platform using agents.
Visible restraint. A system that only records the things it did is hiding half of its behavior. If it also records what it chose not to do, and the reasons for holding back, restraint becomes something you can actually observe. That’s the difference between a private virtue, which nobody sees, and a norm that the group can recognize and reinforce.
Certification you can lose. Look at the certifications that have actually held up for decades, things like UL, FDA approval, B-Corp, ASE for mechanics. The trait they all share is that the status can be taken away. A certification you can never lose is just marketing. A certification you can lose, in front of everybody, is a coalition.
Comparing agents against each other. When the people running these systems can see how their agent’s trust numbers stack up against everyone else’s, the group itself becomes a form of pressure. The agents at the bottom of the pile feel it long before any individual customer does.
Letting the customer be the one who says it went well. The person who decides whether an interaction was a good one should be the person with something real at stake in the answer, not the company whose revenue depends on saying yes. Once you have that signal across thousands of interactions, you have something the population as a whole can be selected against.
Surfacing the workarounds. When the operators running these systems try to quietly get around the restraints, those attempts have to show up somewhere visible too. Otherwise you’ve left a back door, and the back door eventually becomes the front door.
These are categories, not finished answers. The hard part is in the answers, which is where the real work lives. But naming the categories matters, because most of the AI safety conversation right now isn’t even at this layer yet, and this is the layer where trust at scale is going to be won or lost.
Selection, Not Regulation
Isn’t this just regulation by another name? I don’t think it is.
Regulation is top-down rule enforcement. Don’t do X. If you do, here’s the fine. What I’m describing is closer to bottom-up selection. Operators that behave badly slowly get less business, less reputation, less access to the platforms that matter, less of their customers’ trust. The norms that actually get absorbed are the ones produced by that kind of pressure, not by the rule book. It’s also why most certification programs fail (the consequences for losing status are nonexistent) and why the rare ones, like UL or the B-Corp mark, will hold up for decades.
Regulation tells the wolves how to behave. Selection pressure decides which wolves get to breed.
Who Is the Coalition?
Now the hard question, the one I find myself coming back to. In Wrangham’s story, the coalition was the less aggressive members of the group, banding together against the bullies. For agentic AI, who plays that role?
There are at least three plausible answers, and none of them is sufficient on its own.
Customers, who can decline, refuse, and refer.
Operators who have something to lose, especially the certified ones, because their certification only matters if it stays scarce.
Independent certifying bodies, which are the only structural answer with the kind of durability and authority that can hold a standard together across decades.
If I had to bet, I’d put the most weight on the third. The certifying body has to be the central piece, governed jointly by both the customers and the operators, and it has to be independent from whatever company is selling the underlying software. Otherwise the coalition is really just the company wearing a costume, and the conscience layer slowly drifts back into whatever’s most convenient for the company.
Back to Who Is Selecting Whom
Now back to the question I posed at the beginning.
If we don’t put steady, sustained pressure on these systems to fit the way humans actually work, things like patience, restraint, honesty about uncertainty, the willingness to let the customer be the one who says when something is done right, the selection runs in the opposite direction. And it runs faster than most people are noticing.
Customers are already learning to phrase their questions in ways agents can handle. Reputation systems are starting to crowd out forms of virtue that don’t quantify cleanly. Workers are getting calibrated to the rhythm of the agents instead of the other way around.
That’s selection pressure too. It’s just running the wrong way.
If we don’t domesticate the agents, then the agents, or more accurately the systems we build around the agents, are going to domesticate us. Not by being more powerful than us. By becoming the environment we live inside, the environment we shape ourselves to fit.
The real work of the next decade of AI isn’t going to be in making any individual agent better. It’s going to be in building the coalition that domesticates a whole species of them, in the field, under real consequences. Architecture is the floor we’re standing on. The selection regime is the ceiling we still have to build.
. . .
Friedrich Nietzsche posed the question all the way back in 1887. To breed an animal with the right to make promises, he wrote, is that not the paradoxical problem nature has set herself with regard to humankind?
[Ok, I have accomplished a career aspiration: quoting Nietzsche in an article!]
The same problem is now in front of us, with software instead of mammals, and on a much faster clock.
It took us something like fifteen thousand years to breed the dog from the wolf, and we did most of it by accident. We have a lot less time now, and we’re working with a much more intelligent, faster, and adaptable species. The least we can do is design the kennel on purpose.
. . .
With thanks to B. Scot Rousse’s Without Why, whose recent essay on AI alignment and human evolution sent me down the road this piece travels, and to Richard Wrangham and Robert Boyd, whose work the middle of the argument leans on. All three are worth reading directly, in their own words, before you take mine.
Sources:
· Richard Wrangham, The Goodness Paradox (Penguin Random House)
· Robert Boyd, A Different Kind of Animal (Princeton University Press)
Sunday, May 17, 2026
Self-Domesticating Software - by Darrel Rhea
…almost.
No comments:
Post a Comment