Overview
Severity: HIGH | Affected: Multiple LLM Providers | Category: research
Researchers from Carnegie Mellon University's CyLab have published a paper detailing a novel jailbreak technique named 'Unicode Obfuscation Attack' (UOA). This method leverages obscure Unicode characters and non-printable control codes to bypass the safety alignment filters of leading large language models, including those from Anthropic, Google, and OpenAI. The technique involves embedding these characters within prompts, which are ignored by the model's core logic but effectively confuse the input filters designed to detect harmful queries. In their tests, the researchers achieved a success rate of over 85% in generating prohibited content, including hate speech and instructions for illegal activities. The paper has prompted immediate responses from model providers, who are now working to patch their systems against this new class of adversarial attacks.