Anthropic reports the cost of long agent runs
Anthropic reports that Claude Mythos improved an attack on HAWK after 60 hours of work and that each of two highlighted cryptanalysis results cost roughly $100,000 in API use. It says those findings do not affect production cryptography today. Why it matters: Autonomy changes the budget from a per-call number into a cost per verified result. Long runs need a stop condition, checkpoints, and evidence that the result survived independent review.
Try this: For the next long agent job, record a spend cap, deadline, checkpoint, acceptance test, and reviewer before it starts.