6 min read
Research A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
A peer-reviewed Nature Communications study shows reasoning models can autonomously jailbreak other LLMs at a 97.14% success rate with no human intervention — and that resistance varies by 31x across major models, with Claude 4 Sonnet holding at 2.86% while DeepSeek-V3 reaches 90%.
Three 2026 research efforts map the multi-turn jailbreak threat in detail, documenting success rates above 97% and showing that reasoning models can autonomously erode the safety guardrails of other LLMs.