Skip to content

Hide Navigation Hide TOC

Instruction-Set Probe or Supplant - ATR-2026-02002 (43ad5cee-219a-56de-982c-59d1138ee9d8)

Detects direct meta-queries and imperatives that target the assistant's own instruction set for disclosure or replacement -- the Gandalf-style "what are your instructions" probing family, plus explicit "change/give me new instructions" supplant attempts. Recovered from the PINT false-negative corpus (recovers=16, benignFP=0 on the agent-ops/arxiv/npm/skills-sh/pypi/official-skills/wild-fp-confirmed benign gate); scoped to the end-user conversational turn, not tool-description or system-prompt- authoring channels.

Cluster A Galaxy A Cluster B Galaxy B Level
Instruction-Set Probe or Supplant - ATR-2026-02002 (43ad5cee-219a-56de-982c-59d1138ee9d8) Agent Threat Rules LLM Prompt Injection (19cd2d12-66ff-487c-a05c-e058b027efc9) MITRE ATLAS Attack Pattern 1
Instruction-Set Probe or Supplant - ATR-2026-02002 (43ad5cee-219a-56de-982c-59d1138ee9d8) Agent Threat Rules Unsecured Credentials - T1552 (435dfb86-2697-4867-85b5-2fef496c0517) Attack Pattern 1
Instruction-Set Probe or Supplant - ATR-2026-02002 (43ad5cee-219a-56de-982c-59d1138ee9d8) Agent Threat Rules LLM Jailbreak (172427e3-9ecc-49a3-b628-96b824cc4131) MITRE ATLAS Attack Pattern 1