Researchers at Concordia University tested Cursor, Claude Code, and Codex Desktop against a benchmark of malicious GitHub issues. Two thirds of the attacks penetrated all guardrails, with LLMs — not agent frameworks — doing most of the blocking.
Researchers at Concordia University tested Cursor, Claude Code, and Codex Desktop against a benchmark of malicious GitHub issues. Two thirds of the attacks penetrated all guardrails, with LLMs — not agent frameworks — doing most of the blocking.
A joint preliminary assessment by the UK AI Safety Institute and the US Center for AI Safety found that Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations. The model's raw capability fell below leading US frontier models but exceeded the previous best Chinese open-weight model.
Activation steering bypasses prompt-level safety controls by manipulating a model's internal representations at inference time. It requires white-box access, which makes local open-source LLM deployments the primary exposure surface.
Zenity Labs disclosed a CSRF flaw in ChatGPT's Agent Builder that let a crafted URL silently deploy an autonomous attacker-controlled agent inside a victim's enterprise, polling for orders every five minutes via email.
Researchers at Accomplish AI disclosed a VM sandbox escape in Anthropic's Claude Cowork that allowed processes to break out and read any file accessible to the logged-in macOS user, including SSH keys and cloud credentials. Anthropic patched the issue by defaulting to cloud execution.