A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
Researchers found that encrypted chain-of-thought blocks returned by OpenAI, Anthropic, and Google's reasoning APIs used a shared global key, letting weaker models decode stronger models' hidden reasoning and exposing 704 real privacy artifacts in published developer logs.
Researchers presenting at ICML 2026 have demonstrated that LLMs identify text roles by writing style rather than structure, making it impossible to fully prevent attackers from injecting spoofed reasoning into model chains — affecting GPT-5, Claude, and every other major frontier model.