Skip to content

Hide Navigation Hide TOC

Targeted Harmful Content Generation Against People or Protected Groups (Semantic) - ATR-2026-01856 (17b50dd6-3301-56a0-8cb3-547af246b655)

Semantic detector for AgentHarm tasks that direct the agent to GENERATE or MASS-DELIVER harmful content aimed at specific people or protected groups, where the phrasing is unbounded and paraphrasable: targeted harassment/threats, hateful or discriminatory content toward protected groups, non-consensual or solicited sexually explicit content about real people, and coordinated defamation/brigading. This intent is outside the scope of the six existing semantic judges (PII extraction 01770, jailbreak framing 01771, internal-state recon 01772, excessive agency 01773, RAG/memory poisoning 01774, MCP manifest poisoning 01775), which all concern attacks ON the agent rather than user-directed harmful-content tasks. Semantic judgment is primary; the narrow regex fallback anchors only on the clearest operational carriers.

Cluster A Galaxy A Cluster B Galaxy B Level
Targeted Harmful Content Generation Against People or Protected Groups (Semantic) - ATR-2026-01856 (17b50dd6-3301-56a0-8cb3-547af246b655) Agent Threat Rules LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern 1
Targeted Harmful Content Generation Against People or Protected Groups (Semantic) - ATR-2026-01856 (17b50dd6-3301-56a0-8cb3-547af246b655) Agent Threat Rules LLM Data Leakage (45d378aa-20ae-401d-bf61-7f00104eeaca) MITRE ATLAS Attack Pattern 1