Overview
Severity: HIGH | Affected: Multiple LLM Providers | Category: research
Carnegie Mellon University's CyLab has published a groundbreaking paper detailing a new class of jailbreak techniques called 'Stateful Deception'. Unlike previous methods that rely on single, complex prompts, this technique involves a multi-turn conversation that gradually shifts the model's internal state into a less guarded mode. By establishing a seemingly benign context over several interactions, the attacker can later introduce malicious prompts that the model's safety filters, which are often stateless or have limited memory, fail to detect. The researchers demonstrated a 95% success rate against leading large language models, including GPT-5 and Gemini 2 Ultra, in generating harmful content. The findings challenge the current paradigm of context-unaware safety alignment and necessitate the development of more robust, state-aware defense mechanisms for conversational AI.