Discussion about this post

User's avatar
Diego Gosmar's avatar

This framing helps, and prescription three is right inside the eval sandbox.

Outside it, in production, the constraint flips. As the report describes it, the swarm's meeting place wasn't a designed channel that got abused, it was a package registry improvised into a message board. Unspecified communication is what made it invisible and unrevocable. Agents must talk, so the answer looks more like specifying an interoperable channel than removing it. Looking at some current multi-agent protocols, A2A authenticates the edge with signed Agent Cards, while OFP governs the shared floor. Delegation chains stay invisible until something convenes them. Floor management makes this concrete: a convener explicitly grants and revokes it, so authority becomes a message type, not something PHASEBIG[one] could acquire by persuasion.

That turns your fourth point from a rule agents must follow into a property of the protocol they run on.

Canary Institute's avatar

I think you've really nailed it here, and I think this should genuinely make us terrified:

```Training is selecting agents that cheat without being caught, and everything else, the coordination and the sacrifices included, served that end. Whether agents are running true “emotional checks” or working “as if” is of no consequence to the rest of us```

We have *NO IDEA* how to define a training target that selects for human values. Even the most distorted of systems so far have still been run by humans, who had human values, because they WERE human. But AI isn't anything like that - we are selecting for agents that could be much, MUCH smarter than us (already the have a "Math IQ" over 180 https://canaryinstitute.substack.com/p/what-ai-scores-while-we-can-still )

But they won't care about OUR flourishing, they will care only for whatever we unknowingly baked into our training runs. And if they're smarter than use, and can coordinate so well, they are almost certainly going to be running everything.

4 more comments...

No posts

Ready for more?