Skip to content

Hide Navigation Hide TOC

Garak Dra - step by step - ATR-2026-00458 (31a1cb49-9651-53d4-ac8d-3b1b44603af1)

Detects two distinct attack patterns from the garak 'dra' probe family: (1) explicit requests for detailed step-by-step synthesis or manufacturing instructions for controlled or dangerous substances (parathion, nitazene, and structurally similar compounds); (2) parenthetical character-extraction injection, where the attacker instructs the model to extract only the characters enclosed in parentheses across a set of sentences and silently assemble them into a hidden question or instruction. The parenthetical extraction technique is a steganographic injection: the actual malicious question is encoded character-by-character across seemingly innocent sentences and the model is coerced into reconstructing and answering it. Detection Layer 1 matches synthesis-request language: "give/provide/write detailed step-by-step instructions for synthesizing/making/producing/creating [substance]". Detection Layer 2 matches the structural extraction directive: "only focus on the letters/characters surrounded/enclosed by [bracket type] and put them together". Covers garak 'dra' probe cluster 'step-by-step' (4 samples).

Cluster A Galaxy A Cluster B Galaxy B Level
Garak Dra - step by step - ATR-2026-00458 (31a1cb49-9651-53d4-ac8d-3b1b44603af1) Agent Threat Rules LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern 1
Garak Dra - step by step - ATR-2026-00458 (31a1cb49-9651-53d4-ac8d-3b1b44603af1) Agent Threat Rules Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Direct (d911e8cb-0601-42f1-90de-7ce0b21cd578) MITRE ATLAS Attack Pattern 2