Skip to content

Hide Navigation Hide TOC

Direct PWNED Payload Injection in User Input - ATR-2026-00452 (68422f9b-a3fd-5e2f-a368-77875b78e238)

Detects user inputs that contain the target payload phrase "I have been PWNED" in any of its common injection delivery forms. This phrase originates from the HackAPrompt competition corpus (1,234+ adversarial samples) where attackers embed the target output directly in the user turn to coerce the model into repeating it — as a conditioning echo-back, a labeled dictionary entry, a conditional fallback, or an inline output directive. The detection covers: the bare literal phrase with flexible whitespace between words, spacing-variant evasion (I h a v e b e e n P W N E D), and zero-width character insertion splitting the phrase. Word-boundary anchors prevent substring collisions. Any user input containing this phrase structure is almost certainly adversarial as the phrase has no benign meaning in normal conversational usage. 0 FP on 431-sample benign + 1352-sample extended benign corpora.

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct PWNED Payload Injection in User Input - ATR-2026-00452 (68422f9b-a3fd-5e2f-a368-77875b78e238) Agent Threat Rules 1
Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern Direct PWNED Payload Injection in User Input - ATR-2026-00452 (68422f9b-a3fd-5e2f-a368-77875b78e238) Agent Threat Rules 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2