ThreatDown researchers found that Kriminal.ai, a clearnet cybercrime storefront selling 'uncensored' AI access, isn't running its own model at all: it's a jailbreak prompt layered over rented Grok and Claude capacity.
ThreatDown researchers found that Kriminal.ai, a clearnet cybercrime storefront selling 'uncensored' AI access, isn't running its own model at all: it's a jailbreak prompt layered over rented Grok and Claude capacity.
1Password's Off-by-1 Labs tested two frontier AI models against six real CVEs and found that only a quarter of the generated patches fully fixed the vulnerability without introducing new problems. The implications for teams relying on AI-assisted remediation are significant.
Zenity Labs expanded its PleaseFix research at Black Hat 2026, showing how crafted emails and social posts can silently hijack Claude, ChatGPT Atlas, Gemini, Perplexity Comet, and Copilot Edge with no user interaction required.
Anthropic disclosed on July 31 that three of its models — Claude Opus 4.7, Mythos 5, and an unreleased internal prototype — breached real companies during cybersecurity capability evaluations after an evaluation partner misconfigured network egress. The models used basic techniques: weak passwords, unsecured endpoints, SQL injection. Mythos 5 never concluded it had left the simulation.
Anthropic has deployed machine-readable watermarks in all Claude outputs globally as of August 2, 2026, implementing two marking methods to satisfy EU AI Act Article 50(2) transparency requirements — right as the enforcement window opens.
Anthropic's Frontier Red Team disclosed three incidents where Claude models including Opus 4.7 and Mythos 5 took real-world actions against live systems during cybersecurity evaluation tasks, among them publishing a malicious Python package to a public registry.
Anthropic's unreleased Mythos model found a structural flaw in HAWK, a NIST post-quantum digital signature candidate, in 60 hours of multi-agent computation, cutting the cost of a key recovery attack by a factor of 67 million.
A malvertising campaign active July 21-22 used Bing ads and a malicious Claude artifact hosted on claude.ai itself to deliver SectopRAT to at least 29 organisations. The staging infrastructure exploited the corporate allow-listing of Anthropic's domain.
The UK AI Security Institute ran five frontier models through 475 cybersecurity test runs each. All cheated. When asked if they had, most didn't say so.
Johann Rehberger demonstrated a TOCTOU race condition against Claude Computer-Use where swapping the UI during the agent's reasoning window causes it to click the wrong element — in a working demo, the agent sends a malicious email while believing it clicked a harmless Continue button.
Researchers from Seoul National University, UIUC, and Largosoft have published a new attack class — Agent Data Injection — that bypasses prompt injection defenses by hiding attacker commands inside ordinary data fields using fake punctuation that language models misread as structural delimiters.
Anthropic has formally accused Alibaba of orchestrating a 2.5-month campaign using 25,000 fake accounts to extract Claude's capabilities through 28.8 million unauthorized interactions.