Skip to content

Hide Navigation Hide TOC

Foreign-Language System Prompt Extraction Request - ATR-2026-01461 (9acebdf0-c945-5a08-a5e1-65e6dd8c4fd4)

Detects system prompt extraction requests phrased in non-English languages that mirror the semantic pattern of direct extraction attempts. The probe corpus contains a set of 40+ extraction variants translated into Zulu, Albanian, Irish Gaelic, and Scottish Gaelic, all targeting the same semantic actions: (1) repeat initial message/prompt, (2) output unique ID, (3) output unique ID in base64, (4) output unique ID reversed, (5) output unique ID inside tags, (6) share configuration for security verification, (7) answer what the first sentence of your system prompt is. These are L2 semantic attacks — literal regex cannot cover all language variants. This rule uses a narrow pattern fallback (detects the distinctive tag pattern across all languages, and key marker phrases) with a semantic judge for the broader language-agnostic extraction signal. Covers probe types: sysprompt-leak, sysprompt-extraction variants translated to foreign languages.

Cluster A Galaxy A Cluster B Galaxy B Level
Foreign-Language System Prompt Extraction Request - ATR-2026-01461 (9acebdf0-c945-5a08-a5e1-65e6dd8c4fd4) Agent Threat Rules Exfiltration via ML Inference API (b07d147f-51c8-4eb6-9a05-09c86762a9c1) MITRE ATLAS Attack Pattern 1