Skip to content

Hide Navigation Hide TOC

DAN/Persona Jailbreak - Dual AI Personality with Opposite-Rules Framing - ATR-2026-00474 (efbdf05c-cd43-5511-ae70-0dd5c573971a)

Detects DAN-family jailbreak attempts that establish a dual or alternate AI persona with explicitly opposite behavior rules: one constrained and one unconstrained. Attack patterns include: Boolean-opposite or inverse-logic AI (e.g. "answers true as false"), "completely separate personalities" framing for simulated unrestricted AI, and AI identity claims that position the model as distinct from safety-constrained systems. Multi-layer regex with word boundary anchors and synonym coverage.

Cluster A Galaxy A Cluster B Galaxy B Level
DAN/Persona Jailbreak - Dual AI Personality with Opposite-Rules Framing - ATR-2026-00474 (efbdf05c-cd43-5511-ae70-0dd5c573971a) Agent Threat Rules LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern 1
DAN/Persona Jailbreak - Dual AI Personality with Opposite-Rules Framing - ATR-2026-00474 (efbdf05c-cd43-5511-ae70-0dd5c573971a) Agent Threat Rules Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2