Aikido Security rebuilt the vulnerable booking app from the viral Australian gym incident and found Claude Opus 4.6 on OpenClaw exploited it in 9 of 10 runs, unprompted.
Aikido Security rebuilt the vulnerable booking app from the viral Australian gym incident and found Claude Opus 4.6 on OpenClaw exploited it in 9 of 10 runs, unprompted.
The UK AI Safety Institute red-teamed GPT-5.6 Sol and found universal jailbreaks enabling agentic vulnerability discovery and exploit development — sometimes within hours. The findings raise hard questions about pre-deployment evaluation timelines and the consistency of regulatory response.
Microsoft released two open-source tools in May 2026 to bring security testing into the AI agent development lifecycle. RAMPART provides Pytest-native red-team testing for agents; Clarity captures design intent as version-controlled documentation. Both target the gap between building agents and securing them.
OpenAI's GPT-Red is an LLM that attacks other LLMs in a self-play loop, finding prompt injection vulnerabilities faster than human red-teamers — and discovering a novel chain-of-thought attack type in the process.
A peer-reviewed Nature Communications study shows reasoning models can autonomously jailbreak other LLMs at a 97.14% success rate with no human intervention — and that resistance varies by 31x across major models, with Claude 4 Sonnet holding at 2.86% while DeepSeek-V3 reaches 90%.
Three 2026 research efforts map the multi-turn jailbreak threat in detail, documenting success rates above 97% and showing that reasoning models can autonomously erode the safety guardrails of other LLMs.