Overview
Severity: HIGH | Affected: Multiple LLM Providers | Category: research
A new research paper from ETH Zurich has introduced a novel jailbreak technique named 'Semantic Splicing'. This method evades current safety alignment filters on large language models by embedding harmful instructions within seemingly benign, contextually-rich narratives. Instead of using direct commands or simple obfuscation, the technique splices malicious requests into different parts of a long, coherent prompt, which the model reassembles semantically to understand the user's true, harmful intent. The researchers demonstrated successful bypasses against leading models from OpenAI, Google, and Anthropic with a near 95% success rate in their tests for generating misinformation, hate speech, and malicious code. This technique is particularly concerning as it does not rely on specific model weaknesses but exploits the fundamental way models process complex language and context, highlighting the challenge of robust AI safety.