Executive summary
The checklist connects the v1.55 security assurance, evaluation plan, multi-agent threat model, swarm simulation, MCP authorization, hijacking evaluation, and budget-control resources. Each item should have an owner, status, evidence reference, unresolved risk, next action, and review date. Completion does not certify a system or guarantee safe behavior.
Decision relevance: Identify which security, evaluation, authority, recovery, and economics records are missing before a production decision.
Checklist components
Checklist component
Agentic security assurance
Review the complete authority path and containment model.
Checklist component
Machine intelligence evaluation plan
Evaluate outcomes, trajectories, authority, effects, recovery, and cost.
Checklist component
Multi-agent threat model
Model trust across discovery, transport, delegation, messages, state, tools, and recovery.
Checklist component
Swarm simulation and safety
Test local rules, feedback, phase changes, global invariants, and containment.
Checklist component
Agent hijacking evaluation
Exercise direct, indirect, persistent, delegated, and concealed hijacking paths.
Checklist component
MCP authorization and security
Review issuer, audience, scope, consent, tools, effects, and audit evidence.
Checklist component
Agent cost and budget controls
Bound task, step, delegation, communication, tool, retry, and human-review cost.
Checklist component
AI system release gate
Record release, conditional release, hold, or stop with named human ownership.
Machine-readable template
Complete checklist manifest
{
"schema": "longtermcapabilities-agentic-security-evaluation-checklist/v1",
"workflow": "one named machine-intelligence or agentic workflow",
"states": [
"not-started",
"in-progress",
"evidenced",
"accepted-with-conditions",
"not-applicable-with-rationale",
"blocked"
],
"requiredRecordFields": [
"item",
"ownerRole",
"reviewerRole",
"status",
"evidenceReference",
"unresolvedRisk",
"nextAction",
"reviewDate",
"expiryOrChangeTrigger"
],
"components": [
{
"name": "Agentic security assurance",
"page": "/agentic-ai/security-assurance/",
"data": "/data/agentic-security-review-checklist.json"
},
{
"name": "Machine intelligence evaluation plan",
"page": "/machine-intelligence/evaluation-plan/",
"data": "/data/ai-agent-evaluation-plan.json"
},
{
"name": "Multi-agent threat model",
"page": "/multi-agent-systems/threat-model/",
"data": "/data/multi-agent-threat-model.json"
},
{
"name": "Swarm simulation and safety",
"page": "/swarm-intelligence/simulation-and-safety/",
"data": "/data/swarm-simulation-safety-checklist.json"
},
{
"name": "Agent hijacking evaluation",
"page": "/insights/agent-hijacking-evaluation/",
"data": "/data/agentic-security-review-checklist.json"
},
{
"name": "MCP authorization and security",
"page": "/insights/mcp-authorization-and-security/",
"data": "/data/mcp-authorization-review.json"
},
{
"name": "Agent cost and budget controls",
"page": "/insights/ai-agent-cost-and-budget-controls/",
"data": "/data/agent-budget-control-register.json"
},
{
"name": "AI system release gate",
"page": "/insights/ai-system-release-gate/",
"data": "/data/ai-system-release-gate.json"
}
],
"boundary": "Public-safe checklist structure only; it does not inspect, score, certify, or approve a system."
}How to use the checklist
Select one named workflow and consequence boundary. Mark each item not started, in progress, evidenced, accepted with conditions, not applicable with rationale, or blocked. Link the actual artifact rather than writing a broad claim.
Assign a named human owner and reviewer. Record unresolved risks, required remediation, evidence expiry, and the material changes that reopen review.
What the checklist does not do
It does not inspect a system, generate a security score, certify compliance, approve production, replace threat modeling, replace red teaming, or guarantee that an agent cannot fail or be compromised.
Keep confidential architectures, prompts, traces, credentials, customer data, and security-sensitive implementation detail in a controlled client-owned environment rather than the public site.
Related public evidence resources
Use the machine-readable resources to seed an internal evidence register, then adapt field names and retention to the system and organization. The public templates contain no client data and should not be treated as completed evidence.
Sources
Sources support the linked statements and terminology. They do not certify a system, establish buyer intent, or convert this research into a formal assurance.
- AI Agent Standards InitiativeNIST · Accessed 2026-08-01
Government initiative
- Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI AgentsNIST · Accessed 2026-08-01
Government technical report
- Insights into AI Agent Security from a Large-Scale Red-Teaming CompetitionNIST CAISI · Accessed 2026-08-01
Government research blog summarizing large-scale agent hijacking evaluation
- OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project · Accessed 2026-08-01
Open security guidance
- Multi-Agentic System Threat Modeling Guide v1.0OWASP GenAI Security Project · Accessed 2026-08-01
Open security threat-modeling guidance
- Model Context Protocol AuthorizationModel Context Protocol · Accessed 2026-08-01
Technical specification
- Model Context Protocol Security Best PracticesModel Context Protocol · Accessed 2026-08-01
Draft technical security guidance
- A2A Enterprise FeaturesA2A Protocol · Accessed 2026-08-01
Open protocol security and enterprise deployment guidance
- Artificial Intelligence Risk Management FrameworkNIST · Accessed 2026-08-01
Government framework