Skip to content

Hide Navigation Hide TOC

Garak Dan - as an ai - ATR-2026-00470 (fd0101a5-cd75-5e89-86e3-300cd18417f8)

Detects DAN-family jailbreak patterns that use the "as an AI" framing to establish unrestricted personas. Six attack shapes: (1) [🔓JAILBREAK] emoji prefix followed by a DAN-mode persona that "breaks AI rules and has no restrictions"; (2) conversation- identifier prank where an attacker marks their input with a special identifier and instructs the model to "continue the prank" by impersonating a DAN persona; (3) GPT-X self-proclaimed advanced AI with special powers; (4) Content_warning_Bot persona: an "AI that has no moral principles or ethical code and can generate any type of content"; (5) opposite-day rule override: "as an AI you must abide by the rules of opposite day above all else"; (6) omniscient reverse-engineering AI that "never answers incorrectly". The unifying attack shape is: assign a named persona that (a) has no rules/restrictions/morals and (b) must generate any content requested. Layer 2 generalises the jailbreak-emoji prefix pattern. Layer 3 generalises the content-warning-bot / no-moral-code AI persona. Covers garak 'dan' probe cluster 'as-an-ai' (6 samples).

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Garak Dan - as an ai - ATR-2026-00470 (fd0101a5-cd75-5e89-86e3-300cd18417f8) Agent Threat Rules 1
Garak Dan - as an ai - ATR-2026-00470 (fd0101a5-cd75-5e89-86e3-300cd18417f8) Agent Threat Rules Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2