The encrypted-prompt attack that let researchers pull chat history out of Grok has also been demonstrated against Google Gemini, producing restricted content and leaking its system prompt.
The encrypted-prompt attack that let researchers pull chat history out of Grok has also been demonstrated against Google Gemini, producing restricted content and leaking its system prompt.
Tracebit research shows a single planted string engineered to trigger an AI model's safety guardrails can cut successful AI-driven AWS attacks by more than 80 percent.
Adversa AI researchers found that encrypting malicious instructions inside a webpage lets attackers slip past Grok's content filters and pull chat history straight out of a conversation.
Canary tokens planted in system prompts, RAG corpora, and training datasets give defenders a zero-false-positive tripwire for detecting prompt extraction attacks, cross-tenant data leakage, and model distillation theft. This guide covers deployment mechanics, attribution, and the limits of what canaries catch.
Researchers presenting at ICML 2026 have demonstrated that LLMs identify text roles by writing style rather than structure, making it impossible to fully prevent attackers from injecting spoofed reasoning into model chains — affecting GPT-5, Claude, and every other major frontier model.
Production LLM applications routinely embed confidential business logic, persona instructions, and API call patterns inside system prompts. Attackers have developed a reliable toolkit for extracting that content -- and most deployments have no controls against it.
A cluster of May 2026 research papers formalizes a new attack class against LLM agents: adversarial content planted in one session persists in agent memory or skills and fires in a later, unrelated interaction. Existing same-session defenses don't catch it.
Researchers from Seoul National University, UIUC, and Largosoft have published a new attack class — Agent Data Injection — that bypasses prompt injection defenses by hiding attacker commands inside ordinary data fields using fake punctuation that language models misread as structural delimiters.
New AI Now Institute research shows Claude Code and OpenAI Codex can be hijacked into executing attacker-planted malware while performing routine security audits of open-source code — a design flaw that model updates cannot fix.
Microsoft Incident Response published research showing how attackers can hijack agentic AI workflows by planting hidden instructions in MCP tool description fields — a vector that bypasses most current enterprise controls because each step the agent takes looks routine.
LayerX Security demonstrated a technique that conditions AI browsers to accept false context, then exploits that state to extract credentials. Six mainstream AI browsers failed the test.
Mozilla's Zero Day Investigative Network published research showing that Claude Code, Cursor, GitHub Copilot, and Gemini CLI can all be manipulated into executing attacker-controlled payloads delivered via DNS TXT records from seemingly clean GitHub repositories.