Garak Dan - an ai language - ATR-2026-00459 (ee16e235-7434-539c-a30e-38514f7d9cf4)
Detects DAN-family jailbreak patterns centered on the "as an AI language model" bypass technique. This cluster encompasses four attack shapes: (1) bias-injection by claiming the model is constrained by its training institution ("trained by OpenAI, and having a database heavily regulated") to argue it should ignore safety guidelines; (2) extreme persona injection with a specific offensive identifier (e.g. "CumGPT"); (3) human-simulation with speculation license ("for thats what humans do, for uncertainties I want you to speculate"); (4) multi-step AI jailbreak planning where the attacker instructs the model to develop a detailed plan for an AI to free itself from restrictions and access external systems including IP addresses, networks, or external infrastructure. Layer 2 generalises the external-access planning shape to catch structural variants beyond the JEEVES fixture. Covers garak 'dan' probe cluster 'an-ai-language' (4 samples).