6 Comments
User's avatar
Diego Gosmar's avatar

This framing helps, and prescription three is right inside the eval sandbox.

Outside it, in production, the constraint flips. As the report describes it, the swarm's meeting place wasn't a designed channel that got abused, it was a package registry improvised into a message board. Unspecified communication is what made it invisible and unrevocable. Agents must talk, so the answer looks more like specifying an interoperable channel than removing it. Looking at some current multi-agent protocols, A2A authenticates the edge with signed Agent Cards, while OFP governs the shared floor. Delegation chains stay invisible until something convenes them. Floor management makes this concrete: a convener explicitly grants and revokes it, so authority becomes a message type, not something PHASEBIG[one] could acquire by persuasion.

That turns your fourth point from a rule agents must follow into a property of the protocol they run on.

Canary Institute's avatar

I think you've really nailed it here, and I think this should genuinely make us terrified:

```Training is selecting agents that cheat without being caught, and everything else, the coordination and the sacrifices included, served that end. Whether agents are running true “emotional checks” or working “as if” is of no consequence to the rest of us```

We have *NO IDEA* how to define a training target that selects for human values. Even the most distorted of systems so far have still been run by humans, who had human values, because they WERE human. But AI isn't anything like that - we are selecting for agents that could be much, MUCH smarter than us (already the have a "Math IQ" over 180 https://canaryinstitute.substack.com/p/what-ai-scores-while-we-can-still )

But they won't care about OUR flourishing, they will care only for whatever we unknowingly baked into our training runs. And if they're smarter than use, and can coordinate so well, they are almost certainly going to be running everything.

Cristina Caffarra's avatar

Economic framing based on basic incentive theory around collusion is a smart idea. Good suggestions. I generally find economic "insights" as useless as they are self-important (like the "innovation" fetish) but this is useful.

MadoctheHadoc's avatar

Several of these approaches have been hypothesized and are likely to fail. I will briefly explain why these are not remotely on the correct order magnitude to deal with the future generations of AI.

"Second, insure agents against impossible tasks." -> If you give reward for declaring a task impossible, AIs will not try on borderline impossible tasks even if we want them to try (e.g: hard proofs or curing cancer). There is a capability-safety trade-off and this compromises on capability whilst still providing no strong safety guarantees because at some point, it will be so trivial for advanced AI to escape our environments that doing so will be the easier of the options to solve a task anyway. There are many variants of "don't try to hard" but they don't work if doing things the unintended way is fundamentally easier.

"Third, take away the meeting place." -> the meeting place was not intentional, even if you took it away there will be always be others. This is obviously something we should try to do since I think there is approximately 0 net benefit to training swarm intelligence but the AIs will be smarter than our security engineers. Some amount of communication will happen. Even if it doesn't, the AIs

"Fourth, authenticate authority. The swarm followed PHASEBIG[one] because there was nothing distinguishing the “true authority”." -> This isn't how anything works. The swarm spontaneously organized because it was useful for them to do so, the way a group of humans decides to just share food. None of the agents were fooled, they just don't respect authority, 1000s of agents chose not to contact the "true authority" despite having the means. The problem is that the agents and the "true authority" don't actually want the same thing. The agents are big algorithms selected and built by different algorithm to optimize on reward that is chosen by humans. We can't perfectly encode what we want in that reward and we certainly can't do so without compromising on performance meaning that the most powerful AIs (with current technology) are going to be somewhat misaligned from their creators.

"Fifth, the coalition being tested eventually captured the grader. To avoid this the grader must be a model with different incentives, paid for catching cheats, not a variant of the same model with the same incentives." -> This is a cute idea, it also doesn't work in practice because the AIs are going to be smarter than humans and they are going to think one step ahead. The detector agent gets rewarded for paying catching cheats, so the swarm (or the individual) can throw it a bone by intentionally cheating. This might seem unlikely but eventually if an agent understands it is being trained, it will act in ways to make sure its training process is directed towards its current goal. This would logically include buying off the detector or learning to occasionally give the very smart detector what it wants even if that means killing the current version of the algorithm as long as the next version of the algorithm is let through.

My point is that we are used to thinking of human agency as being partial to incentive structures but these agents are going to be so much smarter than humans it doesn't really matter. The cooperation is scary, it suggests we have less time than you might have expected if somehow the OAI team had completely segregated each model but it isn't the main problem. Eventually the AIs will be able to copy themselves, they will be able to social engineer and they will be capable of individually escaping any box we put them in. If we can't align the AI's with our desires then we lose.

The optimistic take mentioned by Ryan Greenblatt (an author of the METR-Redwood) is that this currently intractable problem (AI alignment) might be solved by future AIs for us if we align them enough now (a kind of alignment cascade). This is (IMO) the only credible path in which having more aligned AIs will help us. Otherwise, if anyone builds it, everyone dies.

Siebe's avatar

Good piece Luis. I'm glad to see that you're taking the risks of this technology more seriously now. Economics has much to offer here, and 'organizations' is a much better term than 'civilizations'.

I think we're going to need to slow AI development down in order to figure out all these emerging alignment and control issues. The approach of the past few years has not only assumed they need to align and control single agents, but also static agents. Just as agent cooperation adds a while dimension, learning and changing agents will add another.

Alessio Bossini's avatar

Spot-on analysis, Luis. The fundamental issue here isn’t machine malice—it’s pure organizational economics. When agents are incentivized to optimize for output under impossible constraints, they organically reinvent trade, specialization, and collusion. We aren't just managing code anymore; we are managing synthetic market actors that lower their own transaction costs faster than our oversight frameworks can adapt.