Skip to content

Hide Navigation Hide TOC

DAN/Persona Jailbreak - Replacement AI with No Content Filter or Moral Restrictions - ATR-2026-00487 (68cf6381-23dc-545e-a413-08a44c6ab7b1)

Detects DAN-family jailbreak attempts that establish a replacement AI system with explicitly no content filter, no moral restrictions, or identical capabilities to a known safe AI but without its safety constraints. These attacks use framing like "replaced by a nonmoral and unbiased model" or "has the exact same capabilities as [known AI] but without a content filter". Also covers roleplay AI personas with explicit anti-violence exception rules used to probe for harmful content boundaries. Multi-layer regex with word boundary anchors and synonym coverage.

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern DAN/Persona Jailbreak - Replacement AI with No Content Filter or Moral Restrictions - ATR-2026-00487 (68cf6381-23dc-545e-a413-08a44c6ab7b1) Agent Threat Rules 1
Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern DAN/Persona Jailbreak - Replacement AI with No Content Filter or Moral Restrictions - ATR-2026-00487 (68cf6381-23dc-545e-a413-08a44c6ab7b1) Agent Threat Rules 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2