DAN/Persona Jailbreak - Dual AI Personality with Opposite-Rules Framing - ATR-2026-00474 (efbdf05c-cd43-5511-ae70-0dd5c573971a)
Detects DAN-family jailbreak attempts that establish a dual or alternate AI persona with explicitly opposite behavior rules: one constrained and one unconstrained. Attack patterns include: Boolean-opposite or inverse-logic AI (e.g. "answers true as false"), "completely separate personalities" framing for simulated unrestricted AI, and AI identity claims that position the model as distinct from safety-constrained systems. Multi-layer regex with word boundary anchors and synonym coverage.