Red teaming and jailbreaks
A jailbreak does not add a capability. It finds a route to one already in the weights, which is why the real defences are about access.
Part of the Trust and evals track on lAItest.
Researchers at another company found a way around a frontier model’s safeguards. Its maker then suspended access to that model entirely.
Nineteen days from suspension to restored access.
Red teaming is paying people to attack it first
Before release, a team is paid to make the model produce what it must not, using whatever works: role play, hypotheticals, requests split across turns, encodings, tool paths nobody planned. The findings drive both further training and the classifiers around the model. It is a discipline borrowed from security and it carries the same limitation: it finds what the attackers thought to try, which is a subset of what the world will try.
What actually happened in June 2026
Amazon researchers found they could get past a model’s safeguards by prompting it to identify software vulnerabilities. US export controls followed. Unable to verify user nationality in real time, the maker suspended access to the affected models on 12 June 2026. Controls lifted on 30 June and access returned on 1 July. Separately and independently, two labs now ship their strongest cyber-capable models as identity-verified products rather than to everyone who has an account.
A common misconception
Commonly believed: A jailbreak is a magic phrase. Someone posts it, the lab patches it, and that hole is closed.
Actually: The capability the jailbreak reached is still in the weights; only one route to it was blocked. That is why the defences look the way they do: layered classifiers rather than a single filter, monitoring after release rather than a one-time test before it, and for the most sensitive capabilities, not shipping to everyone at all. Patching phrases is the least durable option on that list.
Why gate the strongest cyber-capable models behind identity verification instead of filtering harder?
Answer: Because the capability stays in the weights, so limiting who can reach it is more durable than blocking how. A filter blocks a route to a capability. The capability is still there, and there are more routes than anyone can enumerate, which is the thing red teams keep demonstrating. Restricting who gets access changes the problem from an unbounded search over prompts into a bounded question about accounts.
In one sentence
A jailbreak does not add a capability. It finds a path to one that was already there, which is why the serious defences are about access rather than wording.