Principal-led architecture for critical systems

Agentic AI evidence resource

A public-safe checklist for the evidence needed before consequential agentic use

Use the checklist to expose missing evidence and assign owners. Do not turn it into a universal score or a substitute for system-specific testing.

Source-linked researchArchitecture guidance with claims and limits visible
Reading time
4 minutes
Reviewed
2026-08-01
Decision relevance
Identify which security, evaluation, authority, recovery, and economics records are missing before a production decision.

Executive summary

The checklist connects the v1.55 security assurance, evaluation plan, multi-agent threat model, swarm simulation, MCP authorization, hijacking evaluation, and budget-control resources. Each item should have an owner, status, evidence reference, unresolved risk, next action, and review date. Completion does not certify a system or guarantee safe behavior.

Decision relevance: Identify which security, evaluation, authority, recovery, and economics records are missing before a production decision.

Checklist components

Machine-readable template

Complete checklist manifest

Open JSON
{
  "schema": "longtermcapabilities-agentic-security-evaluation-checklist/v1",
  "workflow": "one named machine-intelligence or agentic workflow",
  "states": [
    "not-started",
    "in-progress",
    "evidenced",
    "accepted-with-conditions",
    "not-applicable-with-rationale",
    "blocked"
  ],
  "requiredRecordFields": [
    "item",
    "ownerRole",
    "reviewerRole",
    "status",
    "evidenceReference",
    "unresolvedRisk",
    "nextAction",
    "reviewDate",
    "expiryOrChangeTrigger"
  ],
  "components": [
    {
      "name": "Agentic security assurance",
      "page": "/agentic-ai/security-assurance/",
      "data": "/data/agentic-security-review-checklist.json"
    },
    {
      "name": "Machine intelligence evaluation plan",
      "page": "/machine-intelligence/evaluation-plan/",
      "data": "/data/ai-agent-evaluation-plan.json"
    },
    {
      "name": "Multi-agent threat model",
      "page": "/multi-agent-systems/threat-model/",
      "data": "/data/multi-agent-threat-model.json"
    },
    {
      "name": "Swarm simulation and safety",
      "page": "/swarm-intelligence/simulation-and-safety/",
      "data": "/data/swarm-simulation-safety-checklist.json"
    },
    {
      "name": "Agent hijacking evaluation",
      "page": "/insights/agent-hijacking-evaluation/",
      "data": "/data/agentic-security-review-checklist.json"
    },
    {
      "name": "MCP authorization and security",
      "page": "/insights/mcp-authorization-and-security/",
      "data": "/data/mcp-authorization-review.json"
    },
    {
      "name": "Agent cost and budget controls",
      "page": "/insights/ai-agent-cost-and-budget-controls/",
      "data": "/data/agent-budget-control-register.json"
    },
    {
      "name": "AI system release gate",
      "page": "/insights/ai-system-release-gate/",
      "data": "/data/ai-system-release-gate.json"
    }
  ],
  "boundary": "Public-safe checklist structure only; it does not inspect, score, certify, or approve a system."
}

How to use the checklist

Select one named workflow and consequence boundary. Mark each item not started, in progress, evidenced, accepted with conditions, not applicable with rationale, or blocked. Link the actual artifact rather than writing a broad claim.

Assign a named human owner and reviewer. Record unresolved risks, required remediation, evidence expiry, and the material changes that reopen review.

What the checklist does not do

It does not inspect a system, generate a security score, certify compliance, approve production, replace threat modeling, replace red teaming, or guarantee that an agent cannot fail or be compromised.

Keep confidential architectures, prompts, traces, credentials, customer data, and security-sensitive implementation detail in a controlled client-owned environment rather than the public site.

Use the machine-readable resources to seed an internal evidence register, then adapt field names and retention to the system and organization. The public templates contain no client data and should not be treated as completed evidence.

Sources

Sources support the linked statements and terminology. They do not certify a system, establish buyer intent, or convert this research into a formal assurance.

  1. AI Agent Standards InitiativeNIST · Accessed 2026-08-01

    Government initiative

  2. Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI AgentsNIST · Accessed 2026-08-01

    Government technical report

  3. Insights into AI Agent Security from a Large-Scale Red-Teaming CompetitionNIST CAISI · Accessed 2026-08-01

    Government research blog summarizing large-scale agent hijacking evaluation

  4. OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project · Accessed 2026-08-01

    Open security guidance

  5. Multi-Agentic System Threat Modeling Guide v1.0OWASP GenAI Security Project · Accessed 2026-08-01

    Open security threat-modeling guidance

  6. Model Context Protocol AuthorizationModel Context Protocol · Accessed 2026-08-01

    Technical specification

  7. Model Context Protocol Security Best PracticesModel Context Protocol · Accessed 2026-08-01

    Draft technical security guidance

  8. A2A Enterprise FeaturesA2A Protocol · Accessed 2026-08-01

    Open protocol security and enterprise deployment guidance

  9. Artificial Intelligence Risk Management FrameworkNIST · Accessed 2026-08-01

    Government framework

Private local search

Find a service, capability, evidence record, resource, or insight

Press / to open search when focus is not in a form field.

Search runs locally against the public site index.