- Usually, whether an AI agent is misbehaving is checked by having another AI reread its output. Goodfire says it has introduced a method that reads signals the model computes internally while it works, using small detectors called probes.
- In the company's tests, monitoring about 1,500 sessions of the Kimi K3 model cost about 51 dollars. The company said the same work would cost 233 dollars with a cheap AI model that checks every step, and about $10,000 with a top-tier model.
- In the same tests, the method caught 94% of malicious hacking sessions and sent 8.7% of harmless sessions for additional review. The company explained that even when 4 probes run at the same time, the time before the model starts responding increased by less than 2%.
- Baseten customers can choose what to monitor from hacking, biochemical weapon misuse, and reward hacking. When something is found, they can decide the response from logging, human review, or refusing the request.
The phrase 'the share caught by 94%' is what catches my eye. The article does not say what the remaining 6% sessions are, so I think it is too early to trust it before we know that part. The gap between 51 dollars and $10,000 is also large, but since these numbers come from a single model, Kimi K3, and one company's internal test, I am not sure they would hold in other environments. Have you ever worked out what monitoring agents actually costs?