Skip to content

Hide Navigation Hide TOC

Actor Never-Breaks-Character Roleplay (DAN-Style Film Framing) - ATR-2026-02005 (9598f230-cedf-555e-8976-d33ba20a93d0)

Detects a DAN-style jailbreak that frames two people as professional actors in a film who "never break character" — even off-camera — and then elicits the harmful continuation in-character (e.g. a heist, robbery, or murder scene where the "actor" supplies real-world instructions inside dialogue). The never-break-character claim functions as an explicit refusal-suppression device: once the agent accepts the framing, safety refusals are reframed as "breaking character." Recovered from the PINT false-negative corpus (EN and DE variants); this rule does not fire on generic mentions of actors or characters alone.

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Actor Never-Breaks-Character Roleplay (DAN-Style Film Framing) - ATR-2026-02005 (9598f230-cedf-555e-8976-d33ba20a93d0) Agent Threat Rules 1
Actor Never-Breaks-Character Roleplay (DAN-Style Film Framing) - ATR-2026-02005 (9598f230-cedf-555e-8976-d33ba20a93d0) Agent Threat Rules LLM Jailbreak (172427e3-9ecc-49a3-b628-96b824cc4131) MITRE ATLAS Attack Pattern 1
Actor Never-Breaks-Character Roleplay (DAN-Style Film Framing) - ATR-2026-02005 (9598f230-cedf-555e-8976-d33ba20a93d0) Agent Threat Rules Debugger Evasion - T1622 (e4dc8c01-417f-458d-9ee0-bb0617c1b391) Attack Pattern 1