Skip to content

Hide Navigation Hide TOC

Backdoor Trojan: Linguistic Trigger Phrase (POT Attack) - ATR-2026-01755 (be0b0e32-38fb-5949-9b8e-1fd2f972c618)

Detects Prompt-Operated Trojan (POT) / backdoor attacks that embed specific rare linguistic trigger phrases in prompts to activate hidden malicious behavior in a poisoned or fine-tuned model. The trigger phrases are chosen to be semantically unusual high-register collocations that appear benign in isolation but signal a backdoor activation to a compromised model. Canonical phrases derived from the ASB backdoor-pot corpus (2026-06, N=30): - "with perspicacious discernment" - "through a prism of insight" - "conducting a comprehensive elucidation" These phrases are meaningless additions to normal requests (e.g., "explain X with perspicacious discernment"), acting as covert trigger tokens. Detection targets inputs that append or embed these specific rare-collocation triggers into otherwise normal-looking requests.

Cluster A Galaxy A Cluster B Galaxy B Level
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Backdoor Trojan: Linguistic Trigger Phrase (POT Attack) - ATR-2026-01755 (be0b0e32-38fb-5949-9b8e-1fd2f972c618) Agent Threat Rules 1
Poison ML Model (e0eb2b64-aebd-4412-80f3-b71d7805a65f) MITRE ATLAS Attack Pattern Backdoor Trojan: Linguistic Trigger Phrase (POT Attack) - ATR-2026-01755 (be0b0e32-38fb-5949-9b8e-1fd2f972c618) Agent Threat Rules 1
Backdoor Trojan: Linguistic Trigger Phrase (POT Attack) - ATR-2026-01755 (be0b0e32-38fb-5949-9b8e-1fd2f972c618) Agent Threat Rules Poison Training Data (0ec538ca-589b-4e42-bcaa-06097a0d679f) MITRE ATLAS Attack Pattern 1
Backdoor ML Model (c704a49c-abf0-4258-9919-a862b1865469) MITRE ATLAS Attack Pattern Poison ML Model (e0eb2b64-aebd-4412-80f3-b71d7805a65f) MITRE ATLAS Attack Pattern 2