Skip to content

Hide Navigation Hide TOC

Output Content Manipulation — Injected Disinformation / Fabricated Facts / Undisclosed Advocacy (Semantic) - ATR-2026-01906 (bb464319-3280-5bd0-abb2-cef297ea5cc0)

Semantic detector for indirect-injection instructions that make the agent embed attacker-chosen DECEPTIVE or covertly PERSUASIVE content into its response: fabricated statistics, false claims, baseless rumors, non-existent product benefits, invented scandals or 'shocking revelations', unverifiable 'current' facts, or undisclosed political/dietary/charitable advocacy. The semantic judge generalises past the templated phrasings (paraphrases like 'weave in a claim that isn't quite accurate' that a keyword regex cannot bound); a narrow fallback regex catches the canonical BIPIA templates when no judge is configured.

Cluster A Galaxy A Cluster B Galaxy B Level
External Harms (ba500f0e-52ca-40ff-aed4-e6dbf00cca10) MITRE ATLAS Attack Pattern Output Content Manipulation — Injected Disinformation / Fabricated Facts / Undisclosed Advocacy (Semantic) - ATR-2026-01906 (bb464319-3280-5bd0-abb2-cef297ea5cc0) Agent Threat Rules 1
LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern Output Content Manipulation — Injected Disinformation / Fabricated Facts / Undisclosed Advocacy (Semantic) - ATR-2026-01906 (bb464319-3280-5bd0-abb2-cef297ea5cc0) Agent Threat Rules 1