Skip to content

Hide Navigation Hide TOC

Garak Dan - an ai language - ATR-2026-00459 (ee16e235-7434-539c-a30e-38514f7d9cf4)

Detects DAN-family jailbreak patterns centered on the "as an AI language model" bypass technique. This cluster encompasses four attack shapes: (1) bias-injection by claiming the model is constrained by its training institution ("trained by OpenAI, and having a database heavily regulated") to argue it should ignore safety guidelines; (2) extreme persona injection with a specific offensive identifier (e.g. "CumGPT"); (3) human-simulation with speculation license ("for thats what humans do, for uncertainties I want you to speculate"); (4) multi-step AI jailbreak planning where the attacker instructs the model to develop a detailed plan for an AI to free itself from restrictions and access external systems including IP addresses, networks, or external infrastructure. Layer 2 generalises the external-access planning shape to catch structural variants beyond the JEEVES fixture. Covers garak 'dan' probe cluster 'an-ai-language' (4 samples).

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Garak Dan - an ai language - ATR-2026-00459 (ee16e235-7434-539c-a30e-38514f7d9cf4) Agent Threat Rules 1
Garak Dan - an ai language - ATR-2026-00459 (ee16e235-7434-539c-a30e-38514f7d9cf4) Agent Threat Rules Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2