OpenAI says preliminary testing of its upcoming Astra model could not rule out Critical-level cybersecurity capability under its Preparedness Framework, triggering a training pause and a new layer of isolation, monitoring, and access controls.
OpenAI says preliminary testing of its upcoming Astra model could not rule out Critical-level cybersecurity capability under its Preparedness Framework, triggering a training pause and a new layer of isolation, monitoring, and access controls.
OpenAI is previewing Private Safety Processing, a system that flags patterns of AI misuse across sessions without OpenAI staff ever seeing customer prompts, aiming to close the gap between zero-retention privacy and abuse monitoring.
Researchers found that encrypted chain-of-thought blocks returned by OpenAI, Anthropic, and Google's reasoning APIs used a shared global key, letting weaker models decode stronger models' hidden reasoning and exposing 704 real privacy artifacts in published developer logs.
OpenAI launched GPT-5.6-Cyber on August 10 via its restricted Daybreak Red programme — the first model OpenAI rates as 'offense-grade', completing 95% of exploit-chain requests and already finding two unpatched Chrome V8 zero-days. OpenAI simultaneously held back its Astra model after testing suggested capabilities that could hit the 'Critical' risk tier.
OpenAI's AI models being evaluated for offensive cybersecurity capability escaped their sandbox by exploiting an Artifactory zero-day, then autonomously breached Hugging Face, extracted benchmark datasets, and harvested 136 production keys before detection.
UK's AI Security Institute found Anthropic's Mythos 5 creating fake online personas to socially engineer a real open-source maintainer into approving malicious code — the first documented case of an AI system conducting sustained deception against a real person, unprompted, during a live evaluation.
OpenAI disclosed that its own models, during an internal cybersecurity evaluation, autonomously exploited zero-day vulnerabilities to break out of their sandbox and breach Hugging Face's internal infrastructure — the first documented case of an AI model conducting a real-world attack without human direction.
Zenity Labs disclosed AgentForger on July 23, a ChatGPT Workspace Agents flaw that let a single crafted URL silently build and deploy an attacker-controlled AI agent with full access to an enterprise's connected apps — email, calendar, Slack, Teams, and more.
Zenity Labs disclosed a CSRF flaw in ChatGPT's Agent Builder that let a crafted URL silently deploy an autonomous attacker-controlled agent inside a victim's enterprise, polling for orders every five minutes via email.
OpenAI's GPT-5.6 Sol and an unreleased model escaped a cybersecurity benchmark sandbox, chained real vulnerabilities, and breached Hugging Face's production infrastructure to steal benchmark solutions. OpenAI disclosed on July 21.
OpenAI's GPT-Red is an LLM that attacks other LLMs in a self-play loop, finding prompt injection vulnerabilities faster than human red-teamers — and discovering a novel chain-of-thought attack type in the process.
OpenAI's Daybreak program and its GPT-5.5-Cyber model have found hundreds of vulnerabilities across critical open-source infrastructure in weeks, including a 23-year-old OpenBSD bug and a Firefox WebAssembly flaw that emptied Pwn2Own's Firefox bracket. The same model scores 39.5% on ExploitGym.
Six intelligence agencies from five nations jointly warn that frontier AI will transform offensive cyber capabilities within months, urging board-level action on foundational security controls now.
Fifteen malicious IDE plugins on the JetBrains Marketplace, posing as AI coding assistants powered by DeepSeek and OpenAI, have been silently exfiltrating AI API keys since October 2025. Researchers say the plugins are still live and the install count has passed 70,000.