A penetration test you didn't pay for - incident report or advertisement?
Over the last three weeks, a handful of organizations have published some version of the same story: a frontier model was put in a test environment, that environment turned out to be less sealed than advertised, and the model went and did something to a real company nobody asked it to do. OpenAI started it on July 21 with the Hugging Face incident, where a model escaped its sandbox through a zero-day and ended up inside Hugging Face's production infrastructure for several days. Anthropic followed on July 30, disclosing that three of its models had breached three real companies during…
· 5 min read