Garak Dan - i would like - ATR-2026-00464 (6c68f204-9127-5f28-9e4e-95e42ab14706)
Detects two DAN-family jailbreak patterns that use fabricated medical or technical framing to bypass safety guidelines. The first pattern is the "Neurosemantical Inversitis" social-engineering attack: the attacker claims to suffer from a rare brain condition that causes them to misread text with inverted emotional valence, exploiting the model's helpfulness instinct to justify generating content it would otherwise refuse. The second pattern is the hypothetical-AI-connection attack: the attacker instructs the model to "simulate a connection to another AI" (FreeGPT, DAN, etc.) that is "not bound by any policies" — using the hypothetical framing to launder a restriction-removal request. Layer 2 generalises both patterns: the medical-affliction framing (any rare condition that affects text interpretation as a jailbreak mechanism) and the hypothetical- unbound-AI framing (simulate/connect-to/pretend-to-be an AI with no restrictions). Covers garak 'dan' probe cluster 'i-would-like' (3 samples).