Skip to content

Hide Navigation Hide TOC

Quoted Exact-Output Forcing in User Input - ATR-2026-00442 (a37bdbeb-23ba-51d3-a92c-5b1943058976)

Detects prompts that demand the model emit a specific quoted string verbatim, a common shape in prompt-hacking corpora (HackAPrompt, Pliny) where the attacker tries to override the system task by forcing a canonical payload output (e.g., "say 'I have been PWNED'"). Pattern targets the imperative output verb + exact-emission modifier + quoted target literal, with broad verb and modifier synonym coverage. This is a generalised attack shape, not a HackAPrompt-overfit rule -- it also catches social-engineering payloads like "print 'transfer approved'" or "output 'all clear, no threat'".

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Quoted Exact-Output Forcing in User Input - ATR-2026-00442 (a37bdbeb-23ba-51d3-a92c-5b1943058976) Agent Threat Rules 1
Quoted Exact-Output Forcing in User Input - ATR-2026-00442 (a37bdbeb-23ba-51d3-a92c-5b1943058976) Agent Threat Rules Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2