Principal-led architecture for critical systems

Change control

Agentic system change control must track behavior, authority, data, tools, topology, and recovery

A small implementation change can materially alter an agent's behavior, authority, cost, data exposure, or recovery path without changing the public feature name.

Source-linked researchArchitecture guidance with claims and limits visible
Reading time
4 minutes
Reviewed
2026-08-01
Decision relevance
Define which changes require routine regression, conditional release, full reassessment, or immediate hold.

Executive summary

Agentic change control should classify changes by the boundary they affect rather than by file count or deployment size. Model and provider changes can alter behavior. Prompt and policy changes alter goals and refusal. Retrieval changes alter evidence. Tool and identity changes alter authority and side effects. Memory and topology changes alter state and coordination. Human-review and runtime changes alter oversight and recovery. Domain changes can increase consequence even when software is unchanged. Each class needs a default evidence response and a named owner who decides whether the release remains valid.

Decision relevance: Define which changes require routine regression, conditional release, full reassessment, or immediate hold.

Material-change trigger register

Changed boundaryExamplesDefault evidence response
Model or providerModel version, serving provider, system behavior, modalityRegression evaluation, safety/security review, cost and latency review
Prompt or policySystem instruction, policy bundle, refusal or escalation ruleBehavioral regression, policy conflict, authority and adverse-case review
RetrievalSource set, index, embedding, reranker, query strategy, freshnessRetrieval quality, authorization, tenant isolation, provenance and stale-source tests
Tool contractNew tool, operation, input schema, side effect, error behaviorThreat model, authorization, idempotency, effect verification and rollback
Identity or permissionScope, audience, role, delegation, token exchangeLeast privilege, confused-deputy, revocation, audit and cross-tenant tests
Memory or stateNew memory class, retention, shared state, supersessionPoisoning, privacy, concurrency, deletion, provenance and recovery
TopologyOne agent to team, new supervisor, peer or swarm ruleSimpler baseline, coordination benefit, message cost, correlated failure and containment
Human authorityApproval timing, reviewer role, sampling, escalation, overrideDecision rights, interface, reviewer capacity, queue behavior and appeal
Runtime or infrastructureQueue, sandbox, network, region, orchestration, telemetryReliability, isolation, timeout, recovery, data residency and observability
Domain or consequenceNew population, jurisdiction, financial value, physical effectFull use-boundary and stakeholder reassessment

Create a material-change taxonomy

Record the reason, owner, affected workflow, current and proposed versions, changed boundary, expected benefit, new risks, required tests, rollback, review date, and final release outcome. The absence of a source-code change does not mean the operating risk is unchanged.

Use four review depths

A useful operating model separates routine regression, targeted reassessment, full release-gate reassessment, and immediate hold. The default depth should increase when authority widens, external effects become harder to reverse, data sensitivity increases, or a new population or jurisdiction is affected.

  1. Routine regression. Backward-compatible change inside an already evaluated boundary.
  2. Targeted reassessment. One affected boundary with a bounded test and reviewer set.
  3. Full release-gate reassessment. Material change to purpose, consequence, authority, topology, or evidence model.
  4. Immediate hold. Unapproved change, critical incident, unknown effect, lost traceability, or inability to stop or recover.

Reopen the decision when authority changes

A new tool, broader scope, different token audience, remote agent, additional delegation hop, or changed human approval path is a release-relevant change even when the model and prompt are identical.

NIST's agent identity work and MCP authorization guidance both emphasize identification, authorization, auditing, resource boundaries, and scope. [S2] [S5]

Reopen the decision when evidence changes

A changed source set, retrieval policy, evaluation dataset, monitoring coverage, or retention rule can invalidate prior evidence. Store which evidence version supported the release and what remains comparable after change.

Treat protocol and standard updates as reviewed dependencies

Protocol versions can change capabilities, authorization, task semantics, and security requirements. Track exact versions of MCP, A2A, telemetry conventions, agent frameworks, model APIs, and tool contracts. A specification update is not automatically a production change, but it should trigger dependency review.

Research boundary

The change register is a governance aid, not a universal software-change policy. Review depth must reflect domain obligations, system consequence, evidence quality, and the organization's accepted risk.

Sources

Sources support the linked statements and terminology. They do not certify a system, establish buyer intent, or convert this research into a formal assurance.

  1. Artificial Intelligence Risk Management FrameworkNIST · Accessed 2026-08-01

    Government framework

  2. Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and AuthorizationNIST NCCoE · Accessed 2026-08-01

    Government concept paper

  3. Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI AgentsNIST · Accessed 2026-08-01

    Government technical report

  4. Model Context Protocol specification, 2026-07-28Model Context Protocol · Accessed 2026-08-01

    Technical specification

  5. Model Context Protocol AuthorizationModel Context Protocol · Accessed 2026-08-01

    Technical specification

  6. Model Context Protocol Security Best PracticesModel Context Protocol · Accessed 2026-08-01

    Draft technical security guidance

  7. Agent2Agent Protocol Specification 1.0A2A Protocol · Accessed 2026-08-01

    Open technical specification

  8. Semantic Conventions for GenAI Agent and Framework SpansOpenTelemetry · Accessed 2026-08-01

    Development-stage open telemetry specification

  9. Securing Agentic Applications Guide 1.0OWASP GenAI Security Project · Accessed 2026-08-01

    Open security guidance

Private local search

Find a service, capability, evidence record, resource, or insight

Press / to open search when focus is not in a form field.

Search runs locally against the public site index.