Overview
Severity: HIGH | Affected: Multiple LLMs | Category: research
A new paper from the Stanford AI Lab has introduced a novel jailbreak technique called 'Semantic Splicing,' which has proven effective against the safety filters of all major large language models. The attack works by embedding a harmful instruction within a complex, semantically-rich, and benign-looking prompt. The model is tricked into processing the benign context while inadvertently executing the malicious payload, bypassing alignment guardrails. Researchers demonstrated the technique's ability to reliably generate prohibited content, such as detailed phishing emails, malicious code, and sophisticated disinformation, from models like OpenAI's GPT-5 and Anthropic's Claude 4. The findings, published on arXiv, highlight the fundamental challenge of securing models against adversarial inputs and suggest that current safety training methods are insufficient against contextually complex attacks, prompting calls for new defense paradigms.