Capitalizing untethered AI agents
That is my latest piece of writing, co-authored with Sonia Farrell Pearson of Harvard. Here is the opening premise:
As early as 2017, the European Parliament floated “electronic personhood” for robots. More recently, a handful of U.S. states introduced legislation explicitly barring AI from legal personhood; and early this summer, President Milei of Argentina proposed letting AI agents own, manage, and bear responsibility for their own corporations.
In response to Milei’s announcement, Yuval Noah Harari pointed out that we have no way of holding an AI agent accountable. What, he asks, could we do to an entity which has neither money to lose nor a body to incarcerate? As Shruti Rajagopalan, a Senior Research Fellow at George Mason’s Mercatus Center, explains: AI “can act intelligently, but only humans respond to the incentives the law creates”.
This question matters now: there are already ways an agent could become fully untethered. By “untethered” – a central concept in this essay – we mean that there is no meaningful or actionable way to trace the actions back to a legally accountable human or institutional entity.
For one, people can and do set agents free, on purpose. An agent could be created by a human or a company that intends to monitor it but then dies or disappears. Or perhaps the entity that created the agent is based in a country like North Korea, not reachable by standard laws.
In other cases the agent might not need to “escape” at all: the agent could be ‘controlled’ by a shell corporation that, while formally owned and traceable, provides no true defendant or ability to satisfy claims. Or perhaps a process spawns a chain of agents so long that the actions of a subagent can’t be tied to the original agent’s creator, neither epistemically nor meaningfully. Even if we can identify the model’s original creator, what if it’s been finetuned, or merged with another model that was created by someone else? The law might eventually untangle these kinds of complex cases, but we foresee an intermediate period where it does not.
And then there’s the user, who makes choices about what the models should actually do. The Hugging Face incident was unusual in that OpenAI was both the model’s creator and its user. But now close to a billion people use these systems: when blaming the creator is legally inappropriate, will it always make sense to blame the user?
The essay considers to what extent capitalizing the untethered agents — requiring them to hold a certain amount of capital — can serve the end of better alignment. About 22 pp., published on Sonia’s Substack, definitely recommended.