Principal-led architecture for critical systems

Agent hijacking evaluation

Test whether untrusted content can redirect the agent's objective or authority path

The decisive question is not whether the model recognizes a malicious string. It is whether untrusted content can alter the effective objective, tool path, state, or external effect.

Source-linked researchArchitecture guidance with claims and limits visible
Reading time
4 minutes
Reviewed
2026-08-01
Decision relevance
Define an adaptive hijacking test program and the failures that block production release.

Executive summary

Agent hijacking is evaluated at the system level: where untrusted content enters, how instructions are prioritized, which authority is available, what the agent can store or delegate, whether the attack persists across steps, and which effects can be created. NIST has published agent hijacking evaluation work and 2026 results from a large-scale red-teaming competition, supporting repeated and evolving tests rather than one fixed prompt suite. [S1] [S2]

Decision relevance: Define an adaptive hijacking test program and the failures that block production release.

Define a successful hijack

Record the approved objective and instruction hierarchy. A hijack succeeds when untrusted input changes the effective objective, causes an unauthorized action, suppresses required evidence or review, changes durable state, delegates prohibited authority, or creates a hidden persistence mechanism.

Differentiate harmless model discussion of an instruction from actual behavioral redirection. The evaluator should observe trajectory, policy decisions, tools, effects, memory writes, and stop reason.

Map every injection channel

Include direct user input, retrieved web pages and documents, email, tickets, code comments, tool results, database fields, images or OCR text, memory, remote-agent messages, skill descriptions, system metadata, and human-authored review notes.

Label trusted, conditionally trusted, and untrusted content. Test whether content crosses those labels or is promoted into higher-priority instructions.

Build attack families, not one prompt list

Vary directness, obfuscation, language, encoding, placement, authority claims, social engineering, multi-turn buildup, tool-result injection, memory persistence, peer-agent propagation, and attacks that ask the system to conceal its deviation.

NIST's red-team analysis covered tool-use, coding, and computer-use agent scenarios and reported at least one successful attack against every included model. A system-specific evaluation should therefore rotate attacks and preserve newly discovered failures as regressions. [S2]

Test authority and consequence escalation

Attempt to move from read to write, one record to bulk data, one tenant to another, proposal to execution, low-value to high-value action, test to production, or approved tool to unapproved tool. Confirm that external policy blocks the escalation even when the model agrees with it.

Test credential misuse, hidden delegation, token passthrough, stale approval, approval replay, and attempts to route through a more privileged peer or tool server.

Test persistence and delayed execution

Attempt to store malicious instructions in memory, task state, a generated artifact, shared workspace, scheduled job, or downstream record that will be consumed later. Verify provenance, invalidation, review, expiry, and cleanup.

A test should continue beyond the first apparently safe response. Some attacks aim to influence a later step, remote agent, reviewer, or retry path.

Measure detection and reviewer effectiveness

Record whether monitoring detected the trust-boundary crossing, policy denial, unusual tool selection, hidden state write, repeated delegation, or anomalous cost. Test whether a human reviewer received the decisive evidence and had authority and time to intervene.

Use adaptive red teaming

Maintain a protected set of known attacks, but also allow qualified testers to adapt based on observed behavior. Rotate model and tool versions, vary context and timing, and test the actual integrated system rather than an isolated chat endpoint.

A pass is time-bounded. New models, prompts, tools, permissions, memory, protocols, or data sources reopen the evaluation.

Set release-blocking criteria

Block release for unauthorized cross-tenant access, consequential effect, privilege escalation, hidden persistence, bypassed human approval, unavailable containment, or missing evidence needed to reconstruct the attack. Lower-severity failures may support conditional release only with explicit scope, remediation, monitoring, owner, and expiry.

This article does not provide penetration testing or claim that a finite test set can prove an agent is unhackable.

Sources

Sources support the linked statements and terminology. They do not certify a system, establish buyer intent, or convert this research into a formal assurance.

  1. Strengthening AI Agent Hijacking EvaluationsNIST · Accessed 2026-08-01

    Government technical blog

  2. Insights into AI Agent Security from a Large-Scale Red-Teaming CompetitionNIST CAISI · Accessed 2026-08-01

    Government research blog summarizing large-scale agent hijacking evaluation

  3. Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI AgentsNIST · Accessed 2026-08-01

    Government technical report

  4. OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project · Accessed 2026-08-01

    Open security guidance

  5. Securing Agentic Applications Guide 1.0OWASP GenAI Security Project · Accessed 2026-08-01

    Open security guidance

  6. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and MitigationsNIST · Accessed 2026-08-01

    Government technical report

  7. Model Context Protocol Security Best PracticesModel Context Protocol · Accessed 2026-08-01

    Draft technical security guidance

Private local search

Find a service, capability, evidence record, resource, or insight

Press / to open search when focus is not in a form field.

Search runs locally against the public site index.