- Nadella argued in a long post on X that the current situation, in which the only way to check AI judgments is to accept or reject them as they are, must change. He described this as nested black boxes.
- The direction he proposed is a system that isolates models and allows them to be observed, while leaving tamper-proof, human-readable evidence. He also mentioned measures such as disclosing incidents in a timely way, undergoing independent audits, and making data verifiable.
- According to the original text, he said models should be assumed to be compromised from the start and isolated. Authorized people should be able to stop or shut down a model even in the middle of a task, and he added that more advanced models will need standardized advanced isolation techniques.
- This is Nadella's personal argument. The article did not independently verify the effectiveness of the isolation techniques he described. The original text also notes that he refers to AI as super intelligence several times throughout the post.
The article does not say who actually holds the emergency stop authority that Nadella mentioned, and that gap is what stands out to me first. If you are designing a workflow where AI output gets approved at work, I am curious whether you have decided who can halt the model while a task is in progress.