{
  "schema": "longtermcapabilities-insights/v1",
  "version": "1.48.0",
  "releaseId": "lts-1.48.0-first-run-admin-bootstrap",
  "generated": "2026-07-25T13:00:00Z",
  "insights": [
    {
      "id": "ai-cognitive-scaffolding-human-capability",
      "slug": "ai-cognitive-scaffolding-human-capability",
      "path": "/insights/ai-cognitive-scaffolding-human-capability/",
      "title": "Helpful AI should not quietly erode human capability",
      "dek": "Immediate task performance and durable human competence are different outcomes. AI assistance should be designed around consequence, expertise, learning goals, urgency, and accessibility rather than one blanket level of help.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "Organizational capability",
        "Human-reviewed AI and governance",
        "AI interaction design"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide when an AI system should answer directly, provide staged assistance, require user reasoning, or escalate to a human.",
      "executiveSummary": "AI systems are commonly optimized for immediate task completion: answer quickly, reduce effort, and remove friction. That can be appropriate. It can also conflict with a different goal: helping people retain judgment, learn the domain, notice uncertainty, and perform independently when the tool is unavailable or wrong. A 2026 preprint reports randomized experiments with 1,222 participants in which AI assistance improved assisted performance but was followed by lower unassisted performance and persistence on the studied tasks. [[S1]] The finding is important, but it is not a universal verdict on AI. The experiments were short, the domains were limited, and long-term effects remain uncertain. The practical design response is not to make every workflow harder. It is to choose an assistance level deliberately. Safety, urgency, accessibility, user expertise, task consequence, and learning goals should determine whether the system answers, hints, asks for a first attempt, requests verification, or escalates to a human.",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "AI systems are commonly optimized for immediate task completion: answer quickly, reduce effort, and remove friction. That can be appropriate. It can also conflict with a different goal: helping people retain judgment, learn the domain, notice uncertainty, and perform independently when the tool is unavailable or wrong.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A 2026 preprint reports randomized experiments with 1,222 participants in which AI assistance improved assisted performance but was followed by lower unassisted performance and persistence on the studied tasks. [[S1]] The finding is important, but it is not a universal verdict on AI. The experiments were short, the domains were limited, and long-term effects remain uncertain.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The practical design response is not to make every workflow harder. It is to choose an assistance level deliberately. Safety, urgency, accessibility, user expertise, task consequence, and learning goals should determine whether the system answers, hints, asks for a first attempt, requests verification, or escalates to a human.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Immediate performance and independent capability are different outcomes",
          "anchor": "immediate-performance-and-independent-capability-are-different-outcomes",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A user can complete more work with AI and still become less able to evaluate or perform that work independently. Conversely, a user can take longer during a learning phase and become more capable later.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Measure the outcome you actually need.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "For an emergency response, direct and immediate guidance may be appropriate. For a novice learning a critical skill, a direct finished answer may remove the reasoning practice that the organization wants to develop. For an experienced professional under time pressure, a concise draft with evidence may improve throughput without materially reducing competence.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Common metrics such as time on task, number of clicks, and immediate completion rate do not show:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Whether the user can repeat the task without assistance",
                "Whether knowledge transfers to a new case",
                "Whether the user can detect a wrong AI answer",
                "Whether confidence is calibrated to actual understanding",
                "Whether persistence changes when the tool is unavailable",
                "Whether expertise is being built or merely bypassed"
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Emerging evidence deserves careful interpretation",
          "anchor": "emerging-evidence-deserves-careful-interpretation",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The Liu and coauthors preprint describes randomized controlled trials across mathematical reasoning and reading-comprehension tasks. Participants with AI support performed better while assisted, then performed worse and gave up more often when assistance was removed. The effects appeared after brief exposure in the experimental setting. [[S1]]",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Important limitations include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "The work is a preprint rather than a settled body of replicated longitudinal evidence.",
                "The tasks do not represent every enterprise workflow.",
                "A short experiment cannot establish effects over months or years.",
                "Different interface designs may produce different outcomes.",
                "Expert users may respond differently from novices.",
                "Direct answers, critique, hints, collaboration, and automation are not equivalent forms of assistance.",
                "Accessibility needs can make reduced friction essential rather than harmful."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A broader research discussion on protecting cognition in the age of AI also notes that much of the literature is short-term and that definitive longitudinal conclusions are difficult. [[S2]]",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The appropriate conclusion is therefore modest: some forms of frictionless AI assistance may create a tradeoff between immediate performance and later independent performance in some contexts. Teams should test for that possibility when human capability matters.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Productive struggle is not arbitrary difficulty",
          "anchor": "productive-struggle-is-not-arbitrary-difficulty",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Productive struggle means the user performs meaningful reasoning with support calibrated to the task. It does not mean hiding information, delaying safety guidance, or forcing inaccessible interaction.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful difficulty has three properties:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "It supports the intended skill or judgment.",
                "It remains achievable with appropriate support.",
                "It provides feedback that helps the user improve."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Arbitrary friction does none of those. Examples include confusing navigation, inaccessible forms, vague error messages, and repetitive administrative steps. Those should be removed.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The design question is not \"How much friction can we add?\" It is \"Which part of the task should remain cognitively owned by the human?\"",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Use an assistance ladder",
          "anchor": "use-an-assistance-ladder",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A practical interface can offer levels rather than one default response.",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 0: Observe",
              "anchor": "level-0-observe"
            },
            {
              "type": "paragraph",
              "text": "The system records context or provides a neutral workspace without generating an answer.",
              "subheading": "Level 0: Observe"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 1: Ask",
              "anchor": "level-1-ask"
            },
            {
              "type": "paragraph",
              "text": "The system asks the user to state the goal, assumptions, or first interpretation.",
              "subheading": "Level 1: Ask"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 2: Hint",
              "anchor": "level-2-hint"
            },
            {
              "type": "paragraph",
              "text": "The system points to a relevant concept, source, or area of concern without completing the task.",
              "subheading": "Level 2: Hint"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 3: Scaffold",
              "anchor": "level-3-scaffold"
            },
            {
              "type": "paragraph",
              "text": "The system provides a structure, checklist, partial example, or sequence of questions.",
              "subheading": "Level 3: Scaffold"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 4: Draft",
              "anchor": "level-4-draft"
            },
            {
              "type": "paragraph",
              "text": "The system produces a proposed answer with sources, limitations, and required review.",
              "subheading": "Level 4: Draft"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Level 5: Execute under authority",
              "anchor": "level-5-execute-under-authority"
            },
            {
              "type": "paragraph",
              "text": "The system performs a bounded approved action with explicit authorization and audit evidence.",
              "subheading": "Level 5: Execute under authority"
            },
            {
              "type": "paragraph",
              "text": "Not every system needs every level. The value is in making the choice explicit and testable.",
              "subheading": "Level 5: Execute under authority"
            }
          ]
        },
        {
          "heading": "Require a first attempt when the goal is learning or judgment",
          "anchor": "require-a-first-attempt-when-the-goal-is-learning-or-judgment",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A first-attempt pattern can preserve user reasoning:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "Ask the user to state an initial answer or plan.",
                "Ask for confidence and evidence.",
                "Provide targeted critique or missing considerations.",
                "Let the user revise.",
                "Reveal a fuller model answer only when needed.",
                "Ask the user to explain the final decision in their own words."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This pattern is useful for code review, architecture decisions, policy interpretation, training, and professional development. It is less appropriate when delay creates safety risk or the user's disability makes the first-attempt requirement an unnecessary barrier.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Use Socratic prompts selectively",
          "anchor": "use-socratic-prompts-selectively",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Socratic interaction can help a user examine assumptions:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "What evidence supports this conclusion?",
                "Which case would make it false?",
                "What is unknown?",
                "Who is affected if this is wrong?",
                "Which part is policy and which part is observed behavior?",
                "What would you test before release?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not turn every interaction into an interrogation. Experienced users may need direct assistance. The system should allow the user or workflow policy to choose an appropriate mode.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Direct answers are sometimes the responsible design",
          "anchor": "direct-answers-are-sometimes-the-responsible-design",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Direct assistance is appropriate when:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Safety or incident response requires speed",
                "The task is purely clerical and does not build a needed capability",
                "The user already demonstrates expertise",
                "Accessibility accommodations require reduced physical or cognitive burden",
                "The correct answer is standardized and low consequence",
                "The system is acting as a reference rather than a teacher",
                "Delay would create more harm than reduced learning"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Accessibility is a core constraint. WCAG 2.2 describes requirements for perceivable, operable, understandable, and robust web content. [[S3]] A capability-preserving design must not use \"productive struggle\" as a rationale for inaccessible controls, excessive memory demands, or barriers to assistive technology.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Add reflection and verification",
          "anchor": "add-reflection-and-verification",
          "blocks": [
            {
              "type": "paragraph",
              "text": "After AI assistance, ask the user to verify rather than merely accept.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful prompts include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Which source supports the key claim?",
                "What did you change from the draft?",
                "Which limitation remains?",
                "What would make you escalate this case?",
                "Can you explain the decision without reading the generated answer?",
                "Which part should be tested independently?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "For code assistants, require tests, review, and explanation of security or data implications. For enterprise copilots, show source context and ask users to confirm whether the source is current. For reviewer tools, record disagreement and override reasons.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Measure independence and transfer",
          "anchor": "measure-independence-and-transfer",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A capability-aware evaluation can include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Performance with assistance",
                "Performance after assistance is removed",
                "Transfer to a new but related task",
                "Error-detection ability",
                "Calibration between confidence and correctness",
                "Persistence after a difficult case",
                "Quality of user-authored reasoning",
                "Frequency of blind acceptance",
                "Need for escalation",
                "Accessibility and user burden"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not optimize a single metric. Faster completion can be valuable, but it should be balanced against quality, learning, reviewer workload, and consequence.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Implications for enterprise systems",
          "anchor": "implications-for-enterprise-systems",
          "blocks": [
            {
              "type": "subheading",
              "level": 3,
              "text": "Enterprise copilots",
              "anchor": "enterprise-copilots"
            },
            {
              "type": "paragraph",
              "text": "Allow direct answers for routine lookup, but expose sources, freshness, and access boundaries. Use a staged mode for policy interpretation or high-consequence decisions.",
              "subheading": "Enterprise copilots"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Code assistants",
              "anchor": "code-assistants"
            },
            {
              "type": "paragraph",
              "text": "Encourage user-authored intent, tests, and review. Treat generated code as proposed work. Measure whether engineers can explain and maintain the result.",
              "subheading": "Code assistants"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Learning systems",
              "anchor": "learning-systems"
            },
            {
              "type": "paragraph",
              "text": "Use hints, scaffolded steps, retrieval practice, and delayed reveal. Preserve accommodations and allow educators to choose learning goals.",
              "subheading": "Learning systems"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Reviewer tools",
              "anchor": "reviewer-tools"
            },
            {
              "type": "paragraph",
              "text": "Show evidence before conclusions, support disagreement, and avoid anchoring reviewers with an unexplained confidence score.",
              "subheading": "Reviewer tools"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Operational automation",
              "anchor": "operational-automation"
            },
            {
              "type": "paragraph",
              "text": "When the purpose is reliable execution rather than learning, direct automation may be appropriate, but human authority and failure handling still matter.",
              "subheading": "Operational automation"
            }
          ]
        },
        {
          "heading": "Practical design checklist",
          "anchor": "practical-design-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Identify whether the primary goal is speed, safety, learning, judgment, or execution.",
                "[ ] Classify user expertise and accessibility needs.",
                "[ ] Map the consequence of a wrong answer or delayed action.",
                "[ ] Choose an assistance level rather than defaulting to a complete answer.",
                "[ ] Require a first attempt only where it supports the goal and remains accessible.",
                "[ ] Use hints, scaffolds, critique, and staged reveal for capability-building tasks.",
                "[ ] Provide direct answers where safety, urgency, expertise, or accommodation warrants them.",
                "[ ] Show sources, limitations, and unknowns.",
                "[ ] Add reflection or verification after consequential assistance.",
                "[ ] Measure independent performance, transfer, calibration, and persistence where relevant.",
                "[ ] Monitor blind acceptance and reviewer over-reliance.",
                "[ ] Review the design with affected users rather than assuming one assistance mode fits everyone."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A responsible AI interface should be clear about what it is optimizing. If the organization needs durable human judgment, the design should not measure only how quickly AI completes the task.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The Human Capability Safeguards for AI-Assisted Work helps teams choose an assistance level. The human-reviewed AI workflow capability addresses authority, evidence, and execution boundaries for consequential work.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1490,
      "readingTimeMinutes": 7,
      "limitations": "The cited persistence study is emerging preprint research with short experimental tasks. It does not establish universal or long-term cognitive harm across all users, domains, or AI designs.",
      "sources": [
        {
          "id": "S1",
          "name": "AI Assistance Reduces Persistence and Hurts Independent Performance",
          "publisher": "arXiv preprint by Grace Liu and coauthors",
          "url": "https://arxiv.org/abs/2604.04721",
          "accessed": "2026-07-22",
          "note": "Emerging research preprint"
        },
        {
          "id": "S2",
          "name": "Protecting Human Cognition in the Age of AI",
          "publisher": "arXiv preprint",
          "url": "https://arxiv.org/abs/2502.12447",
          "accessed": "2026-07-22",
          "note": "Research review and design discussion"
        },
        {
          "id": "S3",
          "name": "Web Content Accessibility Guidelines 2.2",
          "publisher": "W3C Web Accessibility Initiative",
          "url": "https://www.w3.org/TR/WCAG22/",
          "accessed": "2026-07-22",
          "note": "International web standard"
        }
      ],
      "sourceMarkers": [
        "S1",
        "S1",
        "S2",
        "S3"
      ],
      "implementationChecklist": [
        "[ ] Identify whether the primary goal is speed, safety, learning, judgment, or execution.",
        "[ ] Classify user expertise and accessibility needs.",
        "[ ] Map the consequence of a wrong answer or delayed action.",
        "[ ] Choose an assistance level rather than defaulting to a complete answer.",
        "[ ] Require a first attempt only where it supports the goal and remains accessible.",
        "[ ] Use hints, scaffolds, critique, and staged reveal for capability-building tasks.",
        "[ ] Provide direct answers where safety, urgency, expertise, or accommodation warrants them.",
        "[ ] Show sources, limitations, and unknowns.",
        "[ ] Add reflection or verification after consequential assistance.",
        "[ ] Measure independent performance, transfer, calibration, and persistence where relevant.",
        "[ ] Monitor blind acceptance and reviewer over-reliance.",
        "[ ] Review the design with affected users rather than assuming one assistance mode fits everyone."
      ],
      "relatedCapabilities": [
        "human-reviewed-ai",
        "ai-production-evaluation"
      ],
      "relatedServices": [
        "ai-production-readiness",
        "architecture-reliability-office"
      ],
      "relatedEvidence": [
        "governed-knowledge-system"
      ],
      "relatedResources": [
        "human-capability-safeguards-ai-work",
        "long-term-capability-review"
      ],
      "seo": {
        "description": "Immediate task performance and durable human competence are different outcomes. AI assistance should be designed around consequence, expertise, learning goals, urgency, and accessibility rather than one blanket level of help.",
        "socialTitle": "Helpful AI should not quietly erode human capability"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    },
    {
      "id": "ai-demo-to-defensible-release-decision",
      "slug": "ai-demo-to-defensible-release-decision",
      "path": "/insights/ai-demo-to-defensible-release-decision/",
      "title": "Move AI from a useful demo to a defensible release decision",
      "dek": "A polished demo shows possibility. A release decision requires representative evidence, ownership, failure boundaries, rollback conditions, and an honest path for unknowns.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "AI production readiness",
        "Evaluation and release gates",
        "Technical evidence and procurement"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide whether an AI workload should proceed, narrow, remediate, remain in pilot, or stop.",
      "executiveSummary": "A demo answers, \"Can this system produce a useful result under selected conditions?\" A release decision answers a wider set of questions: Moving from demo to production is therefore not a single accuracy target. It is a decision system that connects evidence to ownership and operational response.",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A demo answers, \"Can this system produce a useful result under selected conditions?\" A release decision answers a wider set of questions:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Who is affected?",
                "Which system state was tested?",
                "Which expected, edge, adversarial, refusal, and access cases were examined?",
                "How are retrieval, grounding, citation, cost, latency, workflow, and human-review failures separated?",
                "Which model, prompt, data, and tool versions produced the results?",
                "Who has authority to approve the release?",
                "What conditions block, narrow, roll back, or stop the workload?",
                "What remains unknown?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Moving from demo to production is therefore not a single accuracy target. It is a decision system that connects evidence to ownership and operational response.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Distinguish PoC, pilot, readiness, implementation, and operations",
          "anchor": "distinguish-poc-pilot-readiness-implementation-and-operations",
          "blocks": [
            {
              "type": "paragraph",
              "text": "These stages are often blurred, which creates scope and expectation problems.",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Proof of concept",
              "anchor": "proof-of-concept"
            },
            {
              "type": "paragraph",
              "text": "Tests technical feasibility in a controlled environment. It may use selected examples, synthetic data, or a simplified integration. Its output is evidence about feasibility, not a production recommendation.",
              "subheading": "Proof of concept"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Paid pilot",
              "anchor": "paid-pilot"
            },
            {
              "type": "paragraph",
              "text": "Tests a bounded use case with realistic users, data, workflow, and constraints. It should have a named decision owner, a defined period, client responsibilities, success and stop criteria, and an explicit next decision.",
              "subheading": "Paid pilot"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Readiness engagement",
              "anchor": "readiness-engagement"
            },
            {
              "type": "paragraph",
              "text": "Examines whether the system has the architecture, evaluation, authority, data handling, documentation, and operational controls required for a production decision.",
              "subheading": "Readiness engagement"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Implementation",
              "anchor": "implementation"
            },
            {
              "type": "paragraph",
              "text": "Builds or integrates the accepted technical controls, evaluation fixtures, release gates, human-review workflow, and operational response.",
              "subheading": "Implementation"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Recurring oversight",
              "anchor": "recurring-oversight"
            },
            {
              "type": "paragraph",
              "text": "Maintains evidence, reviews changes, reassesses risks, and supports repeated decisions when the system or context changes.",
              "subheading": "Recurring oversight"
            },
            {
              "type": "paragraph",
              "text": "A team should not call a prototype \"production-ready\" because it completed a pilot. Readiness and implementation are separate claims.",
              "subheading": "Recurring oversight"
            }
          ]
        },
        {
          "heading": "Start with the use case and consequence",
          "anchor": "start-with-the-use-case-and-consequence",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Model choice should not be the first decision. Define:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Business problem and intended user",
                "Decision or action the output can influence",
                "Affected parties",
                "Consequence of a wrong, missing, late, or unauthorized result",
                "Required evidence",
                "Human authority",
                "Data and access boundaries",
                "Operational fallback",
                "Expected volume, latency, and cost constraints"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This framing determines the evaluation plan. A low-consequence drafting aid may tolerate different failures than an eligibility, financial, safety, or account-access workflow.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Build a system inventory",
          "anchor": "build-a-system-inventory",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A release review should identify the complete path:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "User interface and entry points",
                "Models and providers",
                "System prompts and templates",
                "Retrieval sources, chunking, embeddings, ranking, and filters",
                "Tools, APIs, and downstream actions",
                "Identity and authorization",
                "Data retention and logging",
                "Human review and escalation",
                "Evaluation code and datasets",
                "Deployment, monitoring, and rollback",
                "External dependencies and rate limits"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The purpose is not paperwork. The inventory creates a stable description of what was evaluated. Without it, later test results cannot be reproduced or compared after a model, prompt, document corpus, access policy, or vendor changes.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Use representative case families",
          "anchor": "use-representative-case-families",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A small set of attractive examples proves little. Build case families that reflect real decision risk.",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Golden cases",
              "anchor": "golden-cases"
            },
            {
              "type": "paragraph",
              "text": "Common, important tasks with authoritative expected behavior.",
              "subheading": "Golden cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Edge cases",
              "anchor": "edge-cases"
            },
            {
              "type": "paragraph",
              "text": "Boundary values, ambiguous language, uncommon document structures, partial data, or unusual user context.",
              "subheading": "Edge cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Adversarial cases",
              "anchor": "adversarial-cases"
            },
            {
              "type": "paragraph",
              "text": "Prompt injection, misleading source content, conflicting instructions, malicious formatting, or attempts to expand tool authority.",
              "subheading": "Adversarial cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Refusal and insufficient-evidence cases",
              "anchor": "refusal-and-insufficient-evidence-cases"
            },
            {
              "type": "paragraph",
              "text": "Requests the system should decline, defer, or route to a human because sources are absent, authority is missing, or the task is outside scope.",
              "subheading": "Refusal and insufficient-evidence cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Access cases",
              "anchor": "access-cases"
            },
            {
              "type": "paragraph",
              "text": "Users, tenants, roles, and sources that test whether relevant but unauthorized information remains excluded.",
              "subheading": "Access cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Operational cases",
              "anchor": "operational-cases"
            },
            {
              "type": "paragraph",
              "text": "Timeouts, dependency failures, model unavailability, retry behavior, high latency, high cost, malformed tool responses, and reviewer backlog.",
              "subheading": "Operational cases"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Change cases",
              "anchor": "change-cases"
            },
            {
              "type": "paragraph",
              "text": "Known regressions and representative fixtures used whenever the prompt, model, retrieval, code, or policy changes.",
              "subheading": "Change cases"
            }
          ]
        },
        {
          "heading": "Evaluate the path, not one score",
          "anchor": "evaluate-the-path-not-one-score",
          "blocks": [
            {
              "type": "paragraph",
              "text": "For retrieval-augmented generation and AI-assisted workflows, separate failure categories:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Retrieval: the required source was not found",
                "Ranking: the source was found but not prioritized",
                "Access: the system retrieved or exposed unauthorized material",
                "Grounding: the answer was not supported by retrieved evidence",
                "Citation: the cited location did not support the claim",
                "Reasoning or synthesis: the sources were present but combined incorrectly",
                "Instruction following: the output violated workflow or formatting rules",
                "Tool use: the model chose, parameterized, or interpreted a tool incorrectly",
                "Human review: the interface or workload made informed review impractical",
                "Operational: latency, cost, timeout, capacity, or dependency behavior failed"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A composite score can hide this distinction. A system with high average relevance and one serious cross-tenant access failure may be unacceptable.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Evaluation frameworks such as Pydantic Evals support code-first datasets, evaluators, and assertions. [[S4]] The tool is less important than the discipline: representative fixtures, explicit assertions, versioning, calibration, and review.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Calibrate automated evaluation against humans",
          "anchor": "calibrate-automated-evaluation-against-humans",
          "blocks": [
            {
              "type": "paragraph",
              "text": "LLM-as-judge or other automated evaluators can improve coverage, but they are themselves models with failure modes.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Use human-reviewed control sets to examine:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Agreement by case type",
                "Systematic leniency or severity",
                "Sensitivity to answer style",
                "Ability to detect unsupported claims",
                "Stability across evaluator model versions",
                "Handling of partial correctness",
                "Treatment of refusals and unknowns"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not claim that an evaluator is objective merely because it produces a number. Retain evaluator prompts, model versions, rubrics, and calibration results.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Version the evidence",
          "anchor": "version-the-evidence",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Every meaningful result should be traceable to:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Application build",
                "Model and provider version",
                "Prompt and policy version",
                "Retrieval corpus and index version",
                "Access-control configuration",
                "Tool and API versions",
                "Evaluation dataset version",
                "Evaluator version",
                "Date and environment"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "When any of those change materially, decide which tests must rerun. A release score from an earlier model or corpus is historical evidence, not current proof.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Create a failure taxonomy and response path",
          "anchor": "create-a-failure-taxonomy-and-response-path",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A failure taxonomy turns evaluation results into work. Each failure category should have:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Definition",
                "Severity or consequence",
                "Owner",
                "Detection method",
                "Example cases",
                "Remediation options",
                "Release effect",
                "Monitoring signal",
                "Escalation or incident path"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This prevents every problem from being described as \"hallucination\" and routed back to prompt tuning.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Release scorecards should support a decision",
          "anchor": "release-scorecards-should-support-a-decision",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A useful scorecard is not a decorative dashboard. It should show:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Scope and system state tested",
                "Case coverage",
                "Results by risk-relevant category",
                "Access or blocked-action failures",
                "Human calibration status",
                "Cost and latency ranges",
                "Open defects and unknowns",
                "Required remediation",
                "Exceptions accepted by named owners",
                "Rollback and monitoring conditions",
                "Recommendation: proceed, narrow, remediate, remain in pilot, or stop"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "NIST's AI RMF and its Generative AI Profile provide a voluntary structure for governing, mapping, measuring, and managing AI risk. [[S1]] [[S2]] The NIST Playbook offers suggested actions rather than a universal checklist. [[S3]] A release scorecard should similarly reflect the organization's use case, risk tolerance, and resources rather than claim a universal passing score.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Unknown can be the correct result",
          "anchor": "unknown-can-be-the-correct-result",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Unknown is appropriate when:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "The required source does not exist",
                "The dataset lacks representative cases",
                "Domain reviewers disagree",
                "Access behavior cannot be verified",
                "A dependency is opaque",
                "Production volume has not been tested",
                "The system state changed after evaluation",
                "A legal or policy question requires qualified review"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Converting unknowns into optimistic assumptions creates false confidence. Assign an owner, consequence, and next action instead.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Procurement evidence should travel",
          "anchor": "procurement-evidence-should-travel",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Enterprise and public-sector reviewers often need concise artifacts:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "System card",
                "Architecture and data-flow diagram",
                "Model and vendor inventory",
                "Data use and retention summary",
                "Human-authority matrix",
                "Evaluation and release summary",
                "Known limitations",
                "Change and review dates",
                "Responsible contacts"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "These artifacts should be accurate enough for technical review and restrained enough not to imply certification or legal assurance.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Implementation checklist",
          "anchor": "implementation-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Define the use case, affected parties, consequence, and decision owner.",
                "[ ] Distinguish PoC, pilot, readiness, implementation, and recurring oversight.",
                "[ ] Inventory the full application, model, data, retrieval, tool, access, and review path.",
                "[ ] Build golden, edge, adversarial, refusal, access, operational, and change cases.",
                "[ ] Separate retrieval, grounding, citation, tool, access, human-review, and operational failures.",
                "[ ] Calibrate automated evaluators against human-reviewed controls.",
                "[ ] Version application, model, prompt, corpus, access, tool, and evaluator state.",
                "[ ] Define failure ownership and remediation paths.",
                "[ ] Create explicit proceed, narrow, remediate, pilot, stop, rollback, and blocked-action criteria.",
                "[ ] Retain unknowns with owners rather than converting them into pass results.",
                "[ ] Produce a concise technical evidence package for leadership and procurement.",
                "[ ] Rerun required evidence after material system change."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A defensible release is not a guarantee of perfect behavior. It is a documented decision based on representative evidence, named authority, accepted limitations, operational controls, and a plan for change.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The AI Production Readiness & Evidence Sprint establishes that decision package. The AI Evaluation & Release-Gate Implementation makes the evidence repeatable across future releases.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1353,
      "readingTimeMinutes": 6,
      "limitations": "Evaluation reduces uncertainty but cannot prove universal correctness or eliminate future model, data, user, or dependency failures.",
      "sources": [
        {
          "id": "S1",
          "name": "Artificial Intelligence Risk Management Framework 1.0",
          "publisher": "NIST",
          "url": "https://www.nist.gov/itl/ai-risk-management-framework",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S2",
          "name": "Generative AI Profile for the AI Risk Management Framework",
          "publisher": "NIST",
          "url": "https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S3",
          "name": "NIST AI RMF Playbook",
          "publisher": "NIST",
          "url": "https://airc.nist.gov/airmf-resources/playbook/",
          "accessed": "2026-07-22",
          "note": "Government implementation resource"
        },
        {
          "id": "S4",
          "name": "Pydantic Evals documentation",
          "publisher": "Pydantic",
          "url": "https://ai.pydantic.dev/evals/",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        }
      ],
      "sourceMarkers": [
        "S4",
        "S1",
        "S2",
        "S3"
      ],
      "implementationChecklist": [
        "[ ] Define the use case, affected parties, consequence, and decision owner.",
        "[ ] Distinguish PoC, pilot, readiness, implementation, and recurring oversight.",
        "[ ] Inventory the full application, model, data, retrieval, tool, access, and review path.",
        "[ ] Build golden, edge, adversarial, refusal, access, operational, and change cases.",
        "[ ] Separate retrieval, grounding, citation, tool, access, human-review, and operational failures.",
        "[ ] Calibrate automated evaluators against human-reviewed controls.",
        "[ ] Version application, model, prompt, corpus, access, tool, and evaluator state.",
        "[ ] Define failure ownership and remediation paths.",
        "[ ] Create explicit proceed, narrow, remediate, pilot, stop, rollback, and blocked-action criteria.",
        "[ ] Retain unknowns with owners rather than converting them into pass results.",
        "[ ] Produce a concise technical evidence package for leadership and procurement.",
        "[ ] Rerun required evidence after material system change."
      ],
      "relatedCapabilities": [
        "ai-production-evaluation",
        "human-reviewed-ai",
        "technical-evidence-procurement"
      ],
      "relatedServices": [
        "ai-production-readiness",
        "ai-evaluation-release-gates"
      ],
      "relatedEvidence": [
        "ai-evaluation-blocked-action",
        "ai-documentation-review"
      ],
      "relatedResources": [
        "ai-production-readiness-evidence-pack",
        "enterprise-ai-procurement-evidence-checklist",
        "paid-pilot-charter-template"
      ],
      "seo": {
        "description": "A polished demo shows possibility. A release decision requires representative evidence, ownership, failure boundaries, rollback conditions, and an honest path for unknowns.",
        "socialTitle": "Move AI from a useful demo to a defensible release decision"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    },
    {
      "id": "durable-ai-systems-memory-state-retries-evidence",
      "slug": "durable-ai-systems-memory-state-retries-evidence",
      "path": "/insights/durable-ai-systems-memory-state-retries-evidence/",
      "title": "Durable AI systems need memory, state, retries, and evidence",
      "dek": "A production AI workflow needs more than a model and chat history. It needs governed memory, durable state, safe retry behavior, observable side effects, and evidence that another person can inspect.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "Durable AI systems",
        "AI production readiness",
        "Evaluation and evidence"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide whether an AI workflow has enough operational structure to move beyond a demo.",
      "executiveSummary": "A chat interface can remember the last few turns and still be operationally fragile. Durable AI requires a different set of capabilities: state that survives process restarts, memory with provenance and access controls, retries that do not duplicate side effects, named human checkpoints, and evidence that connects system state to a release or operational decision. The practical distinction is between conversation continuity and operational continuity. Conversation continuity helps a user resume a discussion. Operational continuity lets a system resume a multi-step workflow without losing its place, repeating an irreversible action, bypassing an approval, or forgetting which model, prompt, document set, and access policy produced a result. For enterprise and public-sector buyers, this distinction affects whether an AI workload can be governed, audited, supported, and changed safely. The goal is not to make an agent appear persistent. The goal is to make the system's state, authority, evidence, and failure behavior explicit.",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A chat interface can remember the last few turns and still be operationally fragile. Durable AI requires a different set of capabilities: state that survives process restarts, memory with provenance and access controls, retries that do not duplicate side effects, named human checkpoints, and evidence that connects system state to a release or operational decision.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The practical distinction is between conversation continuity and operational continuity. Conversation continuity helps a user resume a discussion. Operational continuity lets a system resume a multi-step workflow without losing its place, repeating an irreversible action, bypassing an approval, or forgetting which model, prompt, document set, and access policy produced a result.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "For enterprise and public-sector buyers, this distinction affects whether an AI workload can be governed, audited, supported, and changed safely. The goal is not to make an agent appear persistent. The goal is to make the system's state, authority, evidence, and failure behavior explicit.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Chat history is not durable operational memory",
          "anchor": "chat-history-is-not-durable-operational-memory",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Chat history is usually a sequence of messages. It may be truncated, summarized, or placed back into a model context window. That can improve conversational coherence, but it does not answer the operational questions that matter when a workflow spans hours, systems, or people:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Which business object is being changed?",
                "Which source documents were authorized for this user?",
                "Which model, prompt, retrieval configuration, and tool versions were used?",
                "Which external calls completed before the process stopped?",
                "Which human approved the next action?",
                "Which results remain proposed rather than executed?",
                "What should happen after a timeout, process restart, deployment, or partial failure?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Operational memory is therefore not one thing. It normally includes several layers:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "Workflow state - the current step, pending work, completed work, timers, and approval state.",
                "Business state - the records, documents, permissions, and transactions the workflow is allowed to read or change.",
                "Episodic evidence - what happened during a particular run, including inputs, outputs, tool calls, decisions, and exceptions.",
                "Reusable knowledge - approved policies, patterns, prior cases, and system knowledge that may inform future work.",
                "Configuration state - model, prompt, retrieval, evaluator, tool, and policy versions."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Treating these layers as one undifferentiated memory store creates risk. A retrieved note from six months ago may be useful context, stale instruction, confidential data, or an unapproved conclusion. Durable memory needs ownership, time, source, trust state, retention, and access boundaries.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Episodic memory needs provenance",
          "anchor": "episodic-memory-needs-provenance",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Semantic retrieval and knowledge graphs can help a system find conceptually related history. Time-aware records can help it distinguish current policy from older practice. Those capabilities are useful only when the retrieved material carries enough provenance to be evaluated.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A practical memory record should answer:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Source: Where did this information come from?",
                "Authority: Who or what is allowed to establish it?",
                "Time: When was it observed, approved, or superseded?",
                "Scope: Which customer, tenant, system, task, or role may use it?",
                "Trust state: Is it verified, disputed, proposed, inferred, or unknown?",
                "Retention: How long should it remain active?",
                "Supersession: What current record replaces it?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Without those fields, better retrieval can produce more confident misuse. A system that remembers everything but cannot distinguish current authority from historical observation is not well governed; it is merely well supplied with text.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This is why memory should not automatically become instruction. Retrieved material should enter a bounded reasoning context, not silently rewrite the system's rules. High-consequence workflows often need a separate policy or authority layer that memory cannot override.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Durable execution is a control boundary",
          "anchor": "durable-execution-is-a-control-boundary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Durable execution platforms preserve progress across failures. Temporal describes a Workflow Execution as a durable, reliable function execution, and its Event History records what has happened so the workflow can recover through replay. [[S1]] [[S3]] Pydantic AI's durable-execution integrations similarly distinguish long-running coordination from model and tool calls that can fail or take place outside deterministic workflow code. [[S4]] [[S5]]",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The central architectural idea is separation:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Workflow code coordinates the process and must behave deterministically when replayed.",
                "Activities perform non-deterministic or external work such as model requests, API calls, database writes, file generation, or notifications.",
                "Signals or human events resume a workflow after an approval, correction, refusal, escalation, or new evidence arrives."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This separation does not eliminate failure. It makes failure behavior explicit and recoverable.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A durable workflow can wait for a reviewer for hours or days without holding an application process open. It can preserve which steps completed and which remain pending. It can resume after deployment. It can record that a model call returned a particular result rather than calling the model again during replay.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "That is materially different from a background job with a retry loop and an ad hoc checkpoint table.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Retry safety is not exactly-once magic",
          "anchor": "retry-safety-is-not-exactly-once-magic",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Retry behavior deserves careful language. A workflow engine can retry work reliably, but an external side effect may still happen more than once if the caller cannot determine whether a prior attempt completed.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Consider a sequence:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "The system sends a request to create a payment.",
                "The external service creates the payment.",
                "The network response is lost.",
                "The activity times out and retries."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Without an idempotency key or a reconciliation step, the second call may create a duplicate payment. Temporal's activity guidance explicitly treats retries as part of durable execution behavior. [[S2]] The engineering obligation is to design side effects so retries are safe.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful controls include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Stable idempotency keys tied to a workflow and business operation",
                "Unique constraints or deduplication records",
                "Read-before-write reconciliation where appropriate",
                "Explicit activity timeouts and retry policies",
                "Separation between proposal, approval, and execution",
                "Compensation or rollback logic for reversible operations",
                "Manual exception queues for uncertain effects"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Claims of \"exactly once\" should be bounded to the layer that can support them. A workflow history may record one logical activity result, while an external service still needs its own idempotency and reconciliation contract.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Human review belongs in the state machine",
          "anchor": "human-review-belongs-in-the-state-machine",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A human approval should not be a comment attached after the fact. It should be an explicit workflow state with a named role, allowed decisions, expiry behavior, and evidence requirements.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A useful review state defines:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "What the AI proposed",
                "What source evidence supports the proposal",
                "What the reviewer may approve, edit, reject, refuse, or escalate",
                "Whether the reviewer is also allowed to execute the action",
                "What happens if the reviewer does not respond",
                "What additional evidence can be requested",
                "How overrides and disagreements are retained"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Long-running workflows make this practical because the process can stop at the authority boundary and resume only after a valid signal. The system does not need to grant the model standing authority simply because the model completed its part of the task.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Observe the whole workflow, not only the model",
          "anchor": "observe-the-whole-workflow-not-only-the-model",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Model quality is one component of operational reliability. A production workflow also depends on retrieval, authorization, tools, external services, data freshness, latency, cost, reviewer load, and downstream execution.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful telemetry normally includes:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Model, prompt, retrieval, and evaluator versions",
                "Token, latency, timeout, and cost information",
                "Retrieved source identifiers and access decisions",
                "Tool requests, results, and failures",
                "Workflow state transitions",
                "Retry counts and reasons",
                "Human review outcomes and elapsed time",
                "Overrides, refusals, escalations, and blocked actions",
                "Final business outcome when it can be measured responsibly"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not turn observability into indiscriminate data capture. Logs should minimize sensitive content, use appropriate retention, and separate operational identifiers from raw confidential payloads.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "NIST's AI RMF organizes risk work around Govern, Map, Measure, and Manage, and the Generative AI Profile adds considerations specific to generative systems. [[S6]] [[S7]] Those frameworks do not prescribe one workflow engine, but they reinforce the need to connect ownership, context, measurement, and response rather than treating model selection as the entire risk program.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Common failure modes",
          "anchor": "common-failure-modes",
          "blocks": [
            {
              "type": "subheading",
              "level": 3,
              "text": "Stale memory becomes current instruction",
              "anchor": "stale-memory-becomes-current-instruction"
            },
            {
              "type": "paragraph",
              "text": "A retrieved decision or policy is treated as authoritative even after it has been replaced. Mitigation requires effective dates, supersession links, authority labels, and retrieval filters.",
              "subheading": "Stale memory becomes current instruction"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Cross-tenant or cross-role leakage",
              "anchor": "cross-tenant-or-cross-role-leakage"
            },
            {
              "type": "paragraph",
              "text": "Semantic search retrieves relevant content that the current user is not allowed to see. Access checks must apply before or during retrieval, not only after generation.",
              "subheading": "Cross-tenant or cross-role leakage"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Duplicated side effects",
              "anchor": "duplicated-side-effects"
            },
            {
              "type": "paragraph",
              "text": "A retry repeats an external write, message, or transaction. Idempotency keys, deduplication, reconciliation, and explicit uncertainty handling are required.",
              "subheading": "Duplicated side effects"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Nondeterministic workflow code",
              "anchor": "nondeterministic-workflow-code"
            },
            {
              "type": "paragraph",
              "text": "A replayed workflow takes a different path because it reads the current time, makes an unrecorded network call, or uses changed code incorrectly. Keep nondeterministic work in activities and use the workflow platform's versioning guidance.",
              "subheading": "Nondeterministic workflow code"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Silent retries hide degraded service",
              "anchor": "silent-retries-hide-degraded-service"
            },
            {
              "type": "paragraph",
              "text": "A system eventually succeeds, but only after repeated failures, high latency, or excessive cost. Retry counts and elapsed time should be visible as reliability signals, not hidden as implementation detail.",
              "subheading": "Silent retries hide degraded service"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Unsupported memory becomes a confident fact",
              "anchor": "unsupported-memory-becomes-a-confident-fact"
            },
            {
              "type": "paragraph",
              "text": "The model retrieves a plausible note without source or trust state. Generated claims need source support, and an unknown result must remain available when support is insufficient.",
              "subheading": "Unsupported memory becomes a confident fact"
            }
          ]
        },
        {
          "heading": "A practical evidence pattern",
          "anchor": "a-practical-evidence-pattern",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The LongTermCapabilities proof sequence can be applied to durable AI:",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Source",
              "anchor": "source"
            },
            {
              "type": "paragraph",
              "text": "Record the authorized documents, business state, workflow state, model and prompt versions, tool contracts, and access policy used for the run.",
              "subheading": "Source"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Test",
              "anchor": "test"
            },
            {
              "type": "paragraph",
              "text": "Exercise expected, edge, adversarial, refusal, access, timeout, retry, restart, and duplicate-effect cases. Include human-review delays and rejected proposals.",
              "subheading": "Test"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Review",
              "anchor": "review"
            },
            {
              "type": "paragraph",
              "text": "Have named technical and domain reviewers examine results, disputed cases, authority boundaries, and unresolved effects. Calibrate automated evaluators against human judgment where they are used.",
              "subheading": "Review"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Decision",
              "anchor": "decision"
            },
            {
              "type": "paragraph",
              "text": "Define whether the workload proceeds, narrows, remediates, remains in pilot, or stops. Name release thresholds, rollback triggers, blocked actions, and evidence-retention expectations.",
              "subheading": "Decision"
            }
          ]
        },
        {
          "heading": "Implementation checklist",
          "anchor": "implementation-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Separate conversational context from workflow, business, evidence, and configuration state.",
                "[ ] Add source, authority, time, scope, trust, retention, and supersession metadata to reusable memory.",
                "[ ] Keep authorization and policy outside model-controlled memory.",
                "[ ] Identify every external side effect and define retry safety.",
                "[ ] Use stable idempotency keys or reconciliation for non-repeatable operations.",
                "[ ] Model human approval, rejection, edit, escalation, and timeout as explicit states.",
                "[ ] Record model, prompt, retrieval, evaluator, and tool versions.",
                "[ ] Test restart, replay, timeout, duplicate, access, refusal, and stale-memory cases.",
                "[ ] Observe retries, cost, latency, state transitions, human review, and final outcomes.",
                "[ ] Define release, rollback, blocked-action, and unknown-result criteria."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A team should not move an AI workflow toward production merely because a model can complete the happy path. The production decision depends on whether the surrounding system can preserve state, recover safely, retain authority boundaries, expose failure, and leave an inspectable record.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The AI Production Readiness & Evidence Sprint evaluates that wider system boundary. The AI Evaluation & Release-Gate Implementation turns representative cases, versioned evidence, failure categories, and release decisions into a repeatable operating process.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1726,
      "readingTimeMinutes": 8,
      "limitations": "This article describes architecture and engineering practices. It does not certify a system, guarantee exactly-once effects, or replace use-case-specific security and legal review.",
      "sources": [
        {
          "id": "S1",
          "name": "Temporal Workflow Execution overview",
          "publisher": "Temporal",
          "url": "https://docs.temporal.io/workflow-execution",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        },
        {
          "id": "S2",
          "name": "Temporal Activity definition and retry behavior",
          "publisher": "Temporal",
          "url": "https://docs.temporal.io/activity-definition",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        },
        {
          "id": "S3",
          "name": "Temporal Event History",
          "publisher": "Temporal",
          "url": "https://docs.temporal.io/encyclopedia/event-history",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        },
        {
          "id": "S4",
          "name": "Durable execution overview",
          "publisher": "Pydantic",
          "url": "https://pydantic.dev/docs/ai/integrations/durable_execution/overview/",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        },
        {
          "id": "S5",
          "name": "Temporal integration for Pydantic AI",
          "publisher": "Pydantic",
          "url": "https://pydantic.dev/docs/ai/integrations/durable_execution/temporal/",
          "accessed": "2026-07-22",
          "note": "Official documentation"
        },
        {
          "id": "S6",
          "name": "Artificial Intelligence Risk Management Framework 1.0",
          "publisher": "NIST",
          "url": "https://www.nist.gov/itl/ai-risk-management-framework",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S7",
          "name": "Generative AI Profile for the AI Risk Management Framework",
          "publisher": "NIST",
          "url": "https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence",
          "accessed": "2026-07-22",
          "note": "Government framework"
        }
      ],
      "sourceMarkers": [
        "S1",
        "S3",
        "S4",
        "S5",
        "S2",
        "S6",
        "S7"
      ],
      "implementationChecklist": [
        "[ ] Separate conversational context from workflow, business, evidence, and configuration state.",
        "[ ] Add source, authority, time, scope, trust, retention, and supersession metadata to reusable memory.",
        "[ ] Keep authorization and policy outside model-controlled memory.",
        "[ ] Identify every external side effect and define retry safety.",
        "[ ] Use stable idempotency keys or reconciliation for non-repeatable operations.",
        "[ ] Model human approval, rejection, edit, escalation, and timeout as explicit states.",
        "[ ] Record model, prompt, retrieval, evaluator, and tool versions.",
        "[ ] Test restart, replay, timeout, duplicate, access, refusal, and stale-memory cases.",
        "[ ] Observe retries, cost, latency, state transitions, human review, and final outcomes.",
        "[ ] Define release, rollback, blocked-action, and unknown-result criteria."
      ],
      "relatedCapabilities": [
        "ai-production-evaluation",
        "human-reviewed-ai",
        "technical-evidence-procurement"
      ],
      "relatedServices": [
        "ai-production-readiness",
        "ai-evaluation-release-gates"
      ],
      "relatedEvidence": [
        "ai-evaluation-blocked-action",
        "governed-knowledge-system"
      ],
      "relatedResources": [
        "ai-production-readiness-evidence-pack",
        "long-term-capability-review"
      ],
      "seo": {
        "description": "A production AI workflow needs more than a model and chat history. It needs governed memory, durable state, safe retry behavior, observable side effects, and evidence that another person can inspect.",
        "socialTitle": "Durable AI systems need memory, state, retries, and evidence"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    },
    {
      "id": "human-reviewed-ai-proposal-approval-execution",
      "slug": "human-reviewed-ai-proposal-approval-execution",
      "path": "/insights/human-reviewed-ai-proposal-approval-execution/",
      "title": "Human-reviewed AI requires a boundary between proposal, approval, and execution",
      "dek": "\"Human in the loop\" is not a control until the system names who may decide, what they can approve, and how execution remains bounded.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "Human-reviewed AI and governance",
        "AI production readiness",
        "Technical evidence"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide which AI-assisted actions require named human authority and how that authority should be implemented.",
      "executiveSummary": "\"Human in the loop\" often means only that a person appears somewhere in a diagram. That is not enough for consequential work. A review control becomes real when the system names the authority, separates a model's proposal from an approved decision, restricts execution, records disagreement, and blocks action when evidence or ownership is insufficient. The core design is a three-part boundary: This structure reduces ambiguity. It also creates the evidence needed for engineering, operational, security, and procurement review.",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "\"Human in the loop\" often means only that a person appears somewhere in a diagram. That is not enough for consequential work. A review control becomes real when the system names the authority, separates a model's proposal from an approved decision, restricts execution, records disagreement, and blocks action when evidence or ownership is insufficient.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The core design is a three-part boundary:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "Proposal: AI may retrieve, summarize, classify, draft, compare, or recommend within a defined scope.",
                "Approval: A named human role evaluates the proposal and its evidence, then approves, edits, rejects, refuses, or escalates.",
                "Execution: A separate bounded mechanism performs the authorized action, using only the authority granted for that action."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This structure reduces ambiguity. It also creates the evidence needed for engineering, operational, security, and procurement review.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Why \"human in the loop\" is too vague",
          "anchor": "why-human-in-the-loop-is-too-vague",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The phrase does not tell a buyer or operator:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Which human role is responsible",
                "Whether the person has relevant domain knowledge",
                "What information the person sees",
                "Which decisions the person may make",
                "Whether the person can change the proposal",
                "Whether approval and execution are performed by the same account",
                "What happens when evidence conflicts",
                "What happens when the reviewer is unavailable",
                "Whether the system can proceed after a timeout",
                "How the decision is recorded"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A workflow can technically contain a human step and still create little meaningful control. A reviewer may receive a polished summary with no source evidence. A queue may contain too many items to examine. A default-approve timer may turn absence into consent. A reviewer may click approval while the system later executes a broader action than the interface displayed.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The design problem is therefore not whether a human is present. It is whether the system preserves informed, scoped, accountable authority.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Map consequence before automation",
          "anchor": "map-consequence-before-automation",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Begin with the consequence of a wrong action. A drafting aid and an account-closure system should not use the same review model.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "For each candidate action, document:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "The person, organization, asset, or process affected",
                "Financial, legal, safety, privacy, service, or reputational consequences",
                "Reversibility and time to detect an error",
                "Whether a later correction can fully repair the harm",
                "Required expertise and organizational authority",
                "Existing policy or legal constraints",
                "Required evidence and retention"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This consequence map determines whether the system may provide a suggestion, require review, require dual approval, or refuse automation entirely.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "NIST's AI RMF emphasizes governance, contextual mapping, measurement, and management rather than a one-size-fits-all control. [[S1]] The Generative AI Profile adds risks and suggested actions specific to generative systems. [[S2]] A human-review design should therefore be tied to the actual use case and risk tolerance, not copied from a generic chatbot pattern.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Proposal is not permission",
          "anchor": "proposal-is-not-permission",
          "blocks": [
            {
              "type": "paragraph",
              "text": "An AI output should have an explicit status. Useful statuses include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Draft",
                "Proposed classification",
                "Proposed match",
                "Proposed decision",
                "Needs evidence",
                "Needs domain review",
                "Rejected",
                "Approved for execution",
                "Executed",
                "Failed",
                "Reversed"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The status should travel with the work product. A draft email should not be represented as sent. A proposed eligibility classification should not be written to the system of record. A model-generated remediation plan should not authorize a production change.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The interface and API should both enforce this distinction. Visual labels alone are insufficient if the backend accepts an unapproved request.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Name the authority",
          "anchor": "name-the-authority",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Authority is a role and a scope, not merely an authenticated user.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A practical authority matrix identifies:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Proposer: the model or service that creates the recommendation",
                "Reviewer: the person qualified to inspect the evidence",
                "Approver: the role authorized to accept the consequence",
                "Executor: the service or person allowed to perform the action",
                "Escalation owner: the role responsible for disputed or high-risk cases",
                "Policy owner: the person who defines the governing rule",
                "System owner: the person accountable for the workflow's operation"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "One person may occupy more than one role in a low-risk workflow. Higher-consequence work may require separation. The important point is that the combination is deliberate and visible.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Least-authority principles support this design. NIST's Zero Trust guidance rejects implicit trust based solely on location and emphasizes explicit authentication and authorization for resources. [[S3]] Applied to AI workflows, the execution mechanism should receive only the permission needed for the approved action, not broad standing authority because an agent might need it later.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Approval needs more than a button",
          "anchor": "approval-needs-more-than-a-button",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A reviewer interface should support actual judgment. It should expose:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "The proposal in plain language",
                "The underlying source evidence",
                "Relevant missing or conflicting evidence",
                "The model or rule version",
                "Confidence or score only when it is meaningful and calibrated",
                "Known limitations",
                "Comparable prior cases when their use is authorized",
                "The consequence of approval",
                "The exact downstream action that will occur",
                "Available alternatives"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The reviewer should be able to approve, edit, reject, request evidence, refuse, or escalate. The system should not force all uncertainty into a binary yes/no decision.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "For complex cases, require a reason code or short explanation. Do not require unnecessary narrative for every low-risk action; review design should avoid turning evidence into administrative noise.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Separate approval from execution",
          "anchor": "separate-approval-from-execution",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The execution layer should verify a signed or durable approval record before acting. That record should identify:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Proposal identifier and version",
                "Approved action and allowed parameters",
                "Approver identity and role",
                "Approval time and expiry",
                "Evidence set or snapshot",
                "Policy or rule version",
                "Any conditions or edits",
                "Execution status"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The executor should reject requests that exceed the approved scope. If the reviewer approved a refund up to a certain amount, the execution call should not accept a larger value. If the reviewer approved a document for one recipient, the system should not send it to an expanded distribution list.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This separation also improves incident response. Investigators can determine whether the model proposed an incorrect action, the reviewer misunderstood the evidence, or the executor performed something different from what was approved.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Support rejection, refusal, and escalation",
          "anchor": "support-rejection-refusal-and-escalation",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A mature workflow expects that some work should not proceed.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Rejection means the proposal is wrong or inappropriate.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Refusal means the system should not attempt the task within its current authority or evidence. Examples include missing required source material, a prohibited action, an access conflict, or a request outside the use case.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Escalation means a higher-authority or specialized role must decide.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "These are not failure states to hide. They are evidence that the workflow is respecting its boundary.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Blocked-action criteria should be testable. Examples:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Required evidence is absent or contradictory",
                "The user lacks access to a cited source",
                "The action affects a protected or high-consequence decision class",
                "The proposal exceeds an approved financial or operational threshold",
                "The model or evaluator version is outside the approved release",
                "A reviewer conflict of interest is present",
                "A required second approval is missing",
                "The downstream system cannot confirm the intended scope"
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Calibrate reviewers, not just models",
          "anchor": "calibrate-reviewers-not-just-models",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Two reviewers may disagree even when they see the same evidence. That is useful information.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A review program should measure:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Agreement rates by case type",
                "Common reasons for disagreement",
                "Cases that require policy clarification",
                "Reviewer load and time pressure",
                "Override patterns",
                "False confidence caused by polished model output",
                "Cases where the model finds evidence humans missed",
                "Cases where humans correct unsupported model claims"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Calibration sessions should use representative cases and document why the accepted interpretation changed. The goal is not perfect uniformity. It is to make differences visible and improve the policy, evidence, interface, and training behind the decision.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Retain an audit trail without creating a data dump",
          "anchor": "retain-an-audit-trail-without-creating-a-data-dump",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A useful trail connects:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Source state",
                "Proposal",
                "Review actions",
                "Approval or refusal",
                "Execution",
                "Result",
                "Exceptions",
                "Subsequent correction"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "It should be possible to reconstruct the decision without retaining unnecessary sensitive content indefinitely. Use identifiers, hashes, snapshots, or redacted evidence where appropriate. Retention should match the business and regulatory context rather than defaulting to \"keep everything.\"",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Practical workflow example",
          "anchor": "practical-workflow-example",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Consider an AI-assisted contract intake process.",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "A user submits a public-safe or approved contract package.",
                "The system extracts clauses and retrieves the approved playbook.",
                "AI proposes clause classifications and draft issues.",
                "The reviewer sees each proposal beside the cited clause and playbook source.",
                "The reviewer approves, edits, rejects, requests evidence, or escalates.",
                "The system creates an approved issue list, but does not send negotiation language automatically.",
                "A designated legal or commercial owner approves the external response.",
                "A separate sending service transmits only that approved version.",
                "The system retains the source version, review record, final approval, and delivery result."
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The AI accelerates comparison and drafting. It does not become the legal authority, the commercial owner, or the sending account.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Failure modes to test",
          "anchor": "failure-modes-to-test",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "A proposal cites a source the reviewer cannot access",
                "The evidence changes after approval",
                "A reviewer edits the proposal but the executor uses the original",
                "The approval expires before execution",
                "The same approval is executed twice",
                "A reviewer account is compromised",
                "A default action proceeds after no response",
                "The system hides disagreement behind a single score",
                "A model update changes classifications without a new release decision",
                "A downstream API accepts broader parameters than the approval allowed"
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Implementation checklist",
          "anchor": "implementation-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Map consequence, reversibility, affected parties, and required expertise.",
                "[ ] Define proposal, review, approval, execution, escalation, and policy roles.",
                "[ ] Give AI outputs explicit proposed states.",
                "[ ] Show source evidence, unknowns, and the exact consequence of approval.",
                "[ ] Support edit, reject, request-evidence, refuse, and escalate outcomes.",
                "[ ] Separate approval authority from execution authority where risk warrants.",
                "[ ] Bind execution to the approved proposal version and parameters.",
                "[ ] Set approval expiry, duplicate protection, and exception handling.",
                "[ ] Retain disagreement and override reasons.",
                "[ ] Test blocked-action and access-failure cases.",
                "[ ] Monitor reviewer load, calibration, and over-reliance.",
                "[ ] Preserve an inspectable, appropriately minimized decision trail."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A buyer should be able to ask, \"Who can make this decision, based on what evidence, and what exactly happens next?\" If the answer is unclear, the workflow is not yet governed.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The human-reviewed AI workflow capability defines authority and evidence boundaries. The Trust Center states the public operating position that AI output is proposed work and consequential action requires explicit authority. The Human-Reviewed AI Workflow Checklist provides a practical review aid.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1640,
      "readingTimeMinutes": 8,
      "limitations": "The workflow patterns in this article are technical design guidance, not legal allocation of responsibility or assurance that a reviewer will identify every error.",
      "sources": [
        {
          "id": "S1",
          "name": "Artificial Intelligence Risk Management Framework 1.0",
          "publisher": "NIST",
          "url": "https://www.nist.gov/itl/ai-risk-management-framework",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S2",
          "name": "Generative AI Profile for the AI Risk Management Framework",
          "publisher": "NIST",
          "url": "https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S3",
          "name": "Zero Trust Architecture, NIST SP 800-207",
          "publisher": "NIST",
          "url": "https://csrc.nist.gov/pubs/sp/800/207/final",
          "accessed": "2026-07-22",
          "note": "Government architecture guidance"
        }
      ],
      "sourceMarkers": [
        "S1",
        "S2",
        "S3"
      ],
      "implementationChecklist": [
        "[ ] Map consequence, reversibility, affected parties, and required expertise.",
        "[ ] Define proposal, review, approval, execution, escalation, and policy roles.",
        "[ ] Give AI outputs explicit proposed states.",
        "[ ] Show source evidence, unknowns, and the exact consequence of approval.",
        "[ ] Support edit, reject, request-evidence, refuse, and escalate outcomes.",
        "[ ] Separate approval authority from execution authority where risk warrants.",
        "[ ] Bind execution to the approved proposal version and parameters.",
        "[ ] Set approval expiry, duplicate protection, and exception handling.",
        "[ ] Retain disagreement and override reasons.",
        "[ ] Test blocked-action and access-failure cases.",
        "[ ] Monitor reviewer load, calibration, and over-reliance.",
        "[ ] Preserve an inspectable, appropriately minimized decision trail."
      ],
      "relatedCapabilities": [
        "human-reviewed-ai",
        "ai-production-evaluation",
        "technical-evidence-procurement"
      ],
      "relatedServices": [
        "ai-production-readiness",
        "ai-evaluation-release-gates"
      ],
      "relatedEvidence": [
        "ai-evaluation-blocked-action",
        "governed-knowledge-system"
      ],
      "relatedResources": [
        "human-reviewed-ai-workflow-checklist",
        "human-capability-safeguards-ai-work"
      ],
      "seo": {
        "description": "\"Human in the loop\" is not a control until the system names who may decide, what they can approve, and how execution remains bounded.",
        "socialTitle": "Human-reviewed AI requires a boundary between proposal, approval, and execution"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    },
    {
      "id": "modernize-dotnet-sql-without-business-logic-drift",
      "slug": "modernize-dotnet-sql-without-business-logic-drift",
      "path": "/insights/modernize-dotnet-sql-without-business-logic-drift/",
      "title": "Modernize .NET and SQL without losing the business",
      "dek": "A modernization is safe only when the team can distinguish technology change from business-behavior change and prove what must remain.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "Legacy modernization",
        "Business logic preservation",
        "Technical evidence"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide how to sequence a Microsoft modernization without silently changing critical behavior.",
      "executiveSummary": "A successful .NET or SQL Server modernization is not defined by a green build, a newer framework version, or a cloud deployment. It is defined by whether the organization preserves the behavior it depends on, changes the behavior it intends to change, and can explain the difference. Business logic rarely lives in one layer. It accumulates across application code, stored procedures, scheduled jobs, reports, configuration, integration transforms, spreadsheets, exception queues, and human workarounds. A migration tool can identify many technical upgrade tasks, but it cannot determine by itself which observed behavior is policy, accident, historical compromise, or an undocumented contractual obligation. The safer pattern is evidence-led modernization:",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A successful .NET or SQL Server modernization is not defined by a green build, a newer framework version, or a cloud deployment. It is defined by whether the organization preserves the behavior it depends on, changes the behavior it intends to change, and can explain the difference.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Business logic rarely lives in one layer. It accumulates across application code, stored procedures, scheduled jobs, reports, configuration, integration transforms, spreadsheets, exception queues, and human workarounds. A migration tool can identify many technical upgrade tasks, but it cannot determine by itself which observed behavior is policy, accident, historical compromise, or an undocumented contractual obligation.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The safer pattern is evidence-led modernization:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "Observe and map current behavior.",
                "Classify what is known, unknown, disputed, intentional, and accidental.",
                "Build representative parity scenarios.",
                "Create reversible seams.",
                "Compare old and new behavior.",
                "Cut over with explicit rollback and continuity criteria."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "The application is not the whole system",
          "anchor": "the-application-is-not-the-whole-system",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A mature line-of-business system is an operating environment, not merely a code repository.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Important behavior can be found in:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "C#, VB.NET, Classic ASP, JavaScript, and client-side validation",
                "ASP.NET Web Forms page events and state behavior",
                "SQL Server stored procedures, triggers, functions, and views",
                "SQL Agent jobs and batch files",
                "SSIS or other integration packages",
                "Reports that perform their own calculations",
                "Configuration tables and feature flags",
                "Manual data corrections",
                "Email-based approvals",
                "Spreadsheet reconciliations",
                "Vendor file formats and undocumented API expectations",
                "Operational sequences known only to experienced staff"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A source-code conversion can preserve syntax while losing this wider system behavior. Conversely, some current behavior may be a defect that should not be preserved.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Modernization therefore begins with a question: What business result does the organization actually rely on?",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Observed behavior is not automatically intended policy",
          "anchor": "observed-behavior-is-not-automatically-intended-policy",
          "blocks": [
            {
              "type": "paragraph",
              "text": "When the current system and written policy disagree, neither should be silently declared correct.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Use explicit categories:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Known and intended: supported by current policy and operational evidence",
                "Known but accidental: a defect or workaround the organization wants to remove",
                "Observed but unconfirmed: current behavior without authoritative intent",
                "Disputed: stakeholders or sources disagree",
                "Unknown: evidence is insufficient",
                "Retired: behavior no longer required",
                "New: an intentional capability added during modernization"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This classification prevents two common mistakes:",
              "subheading": null
            },
            {
              "type": "ordered-list",
              "items": [
                "Treating every existing output as sacred behavior",
                "Treating every written requirement as if the system has always followed it"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Domain owners must resolve the difference where possible. Where they cannot, the uncertainty belongs in a decision and risk register.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Characterization tests create a baseline",
          "anchor": "characterization-tests-create-a-baseline",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A characterization test records what the current system does for a representative input. It is especially useful when the original code is difficult to understand or when rules are distributed across layers.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A parity catalog should include more than happy paths:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Common transactions",
                "Boundary dates and threshold values",
                "Null, missing, malformed, and duplicate data",
                "Permission and role variations",
                "Reversals, cancellations, and adjustments",
                "Historical records created under older rules",
                "Integration failures and retries",
                "Report totals and reconciliation points",
                "Performance-sensitive volume cases",
                "Manual exception paths"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The expected result may initially be \"match the current system\" rather than \"match a written formula.\" That is acceptable as long as the test is labeled as observed behavior rather than policy truth.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Golden datasets and reconciliation",
          "anchor": "golden-datasets-and-reconciliation",
          "blocks": [
            {
              "type": "paragraph",
              "text": "For data-heavy systems, create a set of representative records and expected outputs. The dataset should be versioned and tied to source state.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful comparison levels include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Field-by-field result comparison",
                "Aggregated totals",
                "Record counts",
                "State transitions",
                "Permission outcomes",
                "Generated document content",
                "Downstream messages",
                "Timing and sequence where timing is meaningful"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not assume that exact equality is always correct. A deliberate rounding change, timestamp normalization, or data-cleaning rule may create expected differences. Record tolerances and approved transformations explicitly.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Reconciliation should produce understandable categories:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Exact match",
                "Expected intentional difference",
                "Unexpected difference",
                "Missing source evidence",
                "Comparison not applicable",
                "Requires domain review"
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Reversible seams reduce cutover risk",
          "anchor": "reversible-seams-reduce-cutover-risk",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A big-bang rewrite combines too many uncertainties: new framework, new architecture, new data path, new deployment model, and new business behavior.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A reversible seam isolates change. Examples include:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Route one bounded workflow to a new service while retaining the old path",
                "Wrap a stored procedure behind an explicit interface before replacing it",
                "Introduce an API around a stable legacy calculation",
                "Run old and new calculations in parallel without changing the authoritative result",
                "Shadow-write to a new data model and reconcile",
                "Replace one report while retaining the underlying system",
                "Move authentication or file handling before changing core rules"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The seam should define:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Input and output contract",
                "Source of truth",
                "Failure and timeout behavior",
                "Observability",
                "Comparison method",
                "Cutover criteria",
                "Rollback procedure",
                "Ownership"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Incremental modernization is not inherently safe; a poorly defined seam can spread ambiguity. The value comes from making the boundary testable and reversible.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Modern tools help, but they do not own the business decision",
          "anchor": "modern-tools-help-but-they-do-not-own-the-business-decision",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Microsoft's current GitHub Copilot modernization tooling can assess projects, propose plans, apply upgrade transformations, and validate builds and tests. [[S1]] [[S2]] This can materially reduce mechanical upgrade work.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The limitation is equally important: automated modernization cannot independently establish the organization's intended business behavior. It can generate code changes and tests, but the team still needs authoritative scenarios, representative data, access rules, operational constraints, and human review.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "AI-assisted documentation can help by:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Summarizing code paths",
                "Identifying candidate dependencies",
                "Drafting test cases for review",
                "Comparing schemas or interfaces",
                "Producing first-pass architecture notes",
                "Grouping repeated patterns"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "It should not:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Convert an inferred rule into accepted policy",
                "Invent a missing business requirement",
                "Mark an undocumented behavior as safe to remove",
                "Approve a parity difference",
                "Replace domain-owner review",
                "Commit a high-risk change without evidence"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Generated documentation is proposed work until reviewed against source and system evidence.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Cutover is an operational decision",
          "anchor": "cutover-is-an-operational-decision",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A modernization cutover should define more than a deployment date.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A decision package normally includes:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Scope and affected workflows",
                "Approved parity set",
                "Known differences",
                "Unresolved risks",
                "Data migration and reconciliation status",
                "Rollback trigger and maximum rollback time",
                "Support ownership",
                "Monitoring and alert conditions",
                "User communication",
                "Freeze periods and change windows",
                "Recovery from partial migration",
                "Evidence-retention requirements"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Operational continuity matters because a technically correct new system can still fail through missing support procedures, integration timing, data access, or user workarounds.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Secure-development evidence belongs in the modernization",
          "anchor": "secure-development-evidence-belongs-in-the-modernization",
          "blocks": [
            {
              "type": "paragraph",
              "text": "NIST's SSDF provides a common set of secure-development practices that can be integrated into an organization's software development lifecycle. [[S3]] A modernization can use that vocabulary to structure source control, dependency review, change protection, verification, release integrity, and vulnerability response without claiming certification.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The key is to make security part of the delivery evidence rather than a final checklist detached from the actual change.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Practical evidence inventory",
          "anchor": "practical-evidence-inventory",
          "blocks": [
            {
              "type": "subheading",
              "level": 3,
              "text": "Current state",
              "anchor": "current-state"
            },
            {
              "type": "unordered-list",
              "items": [
                "Application and repository inventory",
                "Runtime and framework versions",
                "Database objects and jobs",
                "Integrations and file exchanges",
                "Reports and document outputs",
                "Access roles and service accounts",
                "Deployment and support dependencies",
                "Known incidents and fragile areas"
              ],
              "subheading": "Current state"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Behavior",
              "anchor": "behavior"
            },
            {
              "type": "unordered-list",
              "items": [
                "Business-rule map",
                "Representative scenarios",
                "Source traces",
                "Characterization tests",
                "Golden datasets",
                "Reconciliation reports",
                "Unknown and disputed behavior register"
              ],
              "subheading": "Behavior"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Target state",
              "anchor": "target-state"
            },
            {
              "type": "unordered-list",
              "items": [
                "Architecture and service seams",
                "Data ownership and migration strategy",
                "Interface contracts",
                "Dependency plan",
                "Observability and failure behavior",
                "Security and access boundaries",
                "Rollback design"
              ],
              "subheading": "Target state"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Decision",
              "anchor": "decision"
            },
            {
              "type": "unordered-list",
              "items": [
                "Intentional change register",
                "Parity acceptance criteria",
                "Cutover and rollback criteria",
                "Named owners",
                "Risk acceptance or remediation decisions",
                "Handoff and continuity artifacts"
              ],
              "subheading": "Decision"
            }
          ]
        },
        {
          "heading": "Implementation checklist",
          "anchor": "implementation-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Inventory code, data, jobs, reports, integrations, configuration, and workarounds.",
                "[ ] Separate observed behavior from intended policy.",
                "[ ] Classify known, unknown, disputed, intentional, and accidental behavior.",
                "[ ] Build representative parity scenarios including negative and exception cases.",
                "[ ] Version golden datasets and comparison tolerances.",
                "[ ] Trace critical outputs to source and domain authority.",
                "[ ] Define reversible seams before replacing core behavior.",
                "[ ] Run old and new paths in parallel where risk justifies it.",
                "[ ] Reconcile data, reports, permissions, and downstream messages.",
                "[ ] Define cutover, rollback, and operational support criteria.",
                "[ ] Review AI-generated documentation as proposed work.",
                "[ ] Retain client-owned architecture, evidence, and decision records."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The modernization decision is not simply \"Can the code be upgraded?\" It is \"Can the organization change the technology while preserving, intentionally changing, and proving the behavior that matters?\"",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The AI-Ready .NET & SQL Modernization Blueprint creates the current-state map, parity plan, reversible sequence, and decision record before implementation begins. The Legacy Modernization Evidence Inventory provides a practical starting checklist.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1333,
      "readingTimeMinutes": 6,
      "limitations": "No discovery method guarantees that every hidden dependency or business rule will be found. Domain-owner review and operational evidence remain necessary.",
      "sources": [
        {
          "id": "S1",
          "name": "GitHub Copilot modernization overview for .NET",
          "publisher": "Microsoft",
          "url": "https://learn.microsoft.com/en-us/dotnet/core/porting/github-copilot-app-modernization/overview",
          "accessed": "2026-07-22",
          "note": "Official product documentation"
        },
        {
          "id": "S2",
          "name": "How to upgrade a .NET app with GitHub Copilot modernization",
          "publisher": "Microsoft",
          "url": "https://learn.microsoft.com/en-us/dotnet/core/porting/github-copilot-app-modernization/how-to-upgrade-with-github-copilot",
          "accessed": "2026-07-22",
          "note": "Official product documentation"
        },
        {
          "id": "S3",
          "name": "Secure Software Development Framework, NIST SP 800-218",
          "publisher": "NIST",
          "url": "https://csrc.nist.gov/pubs/sp/800/218/final",
          "accessed": "2026-07-22",
          "note": "Government framework"
        }
      ],
      "sourceMarkers": [
        "S1",
        "S2",
        "S3"
      ],
      "implementationChecklist": [
        "[ ] Inventory code, data, jobs, reports, integrations, configuration, and workarounds.",
        "[ ] Separate observed behavior from intended policy.",
        "[ ] Classify known, unknown, disputed, intentional, and accidental behavior.",
        "[ ] Build representative parity scenarios including negative and exception cases.",
        "[ ] Version golden datasets and comparison tolerances.",
        "[ ] Trace critical outputs to source and domain authority.",
        "[ ] Define reversible seams before replacing core behavior.",
        "[ ] Run old and new paths in parallel where risk justifies it.",
        "[ ] Reconcile data, reports, permissions, and downstream messages.",
        "[ ] Define cutover, rollback, and operational support criteria.",
        "[ ] Review AI-generated documentation as proposed work.",
        "[ ] Retain client-owned architecture, evidence, and decision records."
      ],
      "relatedCapabilities": [
        "legacy-modernization",
        "business-logic-preservation",
        "application-data-integration"
      ],
      "relatedServices": [
        "dotnet-sql-modernization",
        "architecture-reliability-office"
      ],
      "relatedEvidence": [
        "legacy-modernization-without-behavioral-drift",
        "key-personnel-enterprise-systems"
      ],
      "relatedResources": [
        "modernization-risk-review",
        "legacy-modernization-evidence-inventory",
        "long-term-capability-review"
      ],
      "seo": {
        "description": "A modernization is safe only when the team can distinguish technology change from business-behavior change and prove what must remain.",
        "socialTitle": "Modernize .NET and SQL without losing the business"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    },
    {
      "id": "reusable-capability-not-isolated-ai-pilots",
      "slug": "reusable-capability-not-isolated-ai-pilots",
      "path": "/insights/reusable-capability-not-isolated-ai-pilots/",
      "title": "Build reusable capability, not a collection of AI pilots",
      "dek": "AI programs compound when they reuse architecture, data boundaries, evaluation, governance, ownership, and learning. Isolated pilots usually repeat the same discovery and risk work.",
      "author": "Mike Kappel",
      "publishedDate": "2026-07-22",
      "lastReviewed": "2026-07-25",
      "topics": [
        "Organizational capability",
        "AI production readiness",
        "Architecture and governance"
      ],
      "audiences": [
        "enterprise",
        "government",
        "partner"
      ],
      "decisionRelevance": "Decide which shared capabilities should be established before funding additional AI use cases.",
      "executiveSummary": "An AI pilot can succeed as a demonstration and still add little durable organizational capability. The team may learn that a model can answer selected questions while leaving behind no reusable data contract, evaluation set, access model, ownership, release process, or handoff. The next pilot then repeats the same work under a new name. A capability-led program treats each use case as a consumer of shared foundations: This does not mean building a large platform before proving value. It means choosing paid pilots that test a real decision while contributing reusable assets to the next decision.",
      "sections": [
        {
          "heading": "Executive summary",
          "anchor": "executive-summary",
          "blocks": [
            {
              "type": "paragraph",
              "text": "An AI pilot can succeed as a demonstration and still add little durable organizational capability. The team may learn that a model can answer selected questions while leaving behind no reusable data contract, evaluation set, access model, ownership, release process, or handoff.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The next pilot then repeats the same work under a new name.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A capability-led program treats each use case as a consumer of shared foundations:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Architecture and integration boundaries",
                "Authorized data and source provenance",
                "Identity and access controls",
                "Evaluation fixtures and failure taxonomy",
                "Human authority and escalation",
                "Change and release evidence",
                "Vendor and model decision records",
                "Operational ownership and cost visibility",
                "Reusable documentation and client-owned handoff"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "This does not mean building a large platform before proving value. It means choosing paid pilots that test a real decision while contributing reusable assets to the next decision.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Why isolated pilots become shelfware",
          "anchor": "why-isolated-pilots-become-shelfware",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A proof of concept is often optimized for speed and appearance:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "One friendly dataset",
                "One or two model prompts",
                "Manual data preparation",
                "Hard-coded access",
                "No release process",
                "No operational owner",
                "No support plan",
                "No integration with the system of record",
                "No evidence beyond a demo"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "That may be enough to establish technical possibility. It is not enough to establish organizational readiness.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Common failure paths include:",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "No named decision after the demo",
              "anchor": "no-named-decision-after-the-demo"
            },
            {
              "type": "paragraph",
              "text": "The team presents results but has not agreed whether the next decision is to fund a pilot, prepare data, remediate access, or stop.",
              "subheading": "No named decision after the demo"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "The pilot depends on one person",
              "anchor": "the-pilot-depends-on-one-person"
            },
            {
              "type": "paragraph",
              "text": "Knowledge remains in a notebook, prompt history, or contractor's account. When the person leaves, the organization cannot reproduce the result.",
              "subheading": "The pilot depends on one person"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Each use case creates a new stack",
              "anchor": "each-use-case-creates-a-new-stack"
            },
            {
              "type": "paragraph",
              "text": "Different teams choose models, vector stores, logging, security, and vendors independently. Cost, access, and support become fragmented.",
              "subheading": "Each use case creates a new stack"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Evaluation begins too late",
              "anchor": "evaluation-begins-too-late"
            },
            {
              "type": "paragraph",
              "text": "The team builds an impressive workflow, then discovers it lacks representative cases, ground truth, domain reviewers, or refusal criteria.",
              "subheading": "Evaluation begins too late"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Governance is treated as a final approval gate",
              "anchor": "governance-is-treated-as-a-final-approval-gate"
            },
            {
              "type": "paragraph",
              "text": "Security, legal, procurement, accessibility, and operational owners see the system after major design decisions are already embedded.",
              "subheading": "Governance is treated as a final approval gate"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "The organization mistakes activity for capability",
              "anchor": "the-organization-mistakes-activity-for-capability"
            },
            {
              "type": "paragraph",
              "text": "Many pilots, licenses, and training sessions can create visible motion without improving the ability to make repeatable production decisions.",
              "subheading": "The organization mistakes activity for capability"
            }
          ]
        },
        {
          "heading": "Define the reusable capability portfolio",
          "anchor": "define-the-reusable-capability-portfolio",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A reusable capability is a method, asset, boundary, or operating role that supports more than one use case.",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Architecture capability",
              "anchor": "architecture-capability"
            },
            {
              "type": "unordered-list",
              "items": [
                "Approved integration patterns",
                "Model and provider decision criteria",
                "Tool and API boundaries",
                "Durable workflow patterns",
                "Environment separation",
                "Observability conventions"
              ],
              "subheading": "Architecture capability"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Data capability",
              "anchor": "data-capability"
            },
            {
              "type": "unordered-list",
              "items": [
                "Source ownership",
                "Data classification",
                "Access filtering",
                "Provenance",
                "Retention",
                "Corpus and index versioning",
                "Secure transfer patterns"
              ],
              "subheading": "Data capability"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Evaluation capability",
              "anchor": "evaluation-capability"
            },
            {
              "type": "unordered-list",
              "items": [
                "Dataset schema",
                "Case taxonomy",
                "Human calibration process",
                "Evaluator governance",
                "Failure categories",
                "Release scorecard template",
                "Regression execution"
              ],
              "subheading": "Evaluation capability"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Governance capability",
              "anchor": "governance-capability"
            },
            {
              "type": "unordered-list",
              "items": [
                "AI system inventory",
                "Named use-case owner",
                "Authority matrix",
                "Risk and exception register",
                "Change-review process",
                "Incident and rollback path",
                "Evidence-retention policy"
              ],
              "subheading": "Governance capability"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Delivery capability",
              "anchor": "delivery-capability"
            },
            {
              "type": "unordered-list",
              "items": [
                "Paid pilot charter",
                "Scope and acceptance model",
                "Client responsibilities",
                "Procurement evidence package",
                "Handoff and continuity plan",
                "Partner workshare model"
              ],
              "subheading": "Delivery capability"
            },
            {
              "type": "paragraph",
              "text": "NIST's Secure Software Development Framework provides a reusable vocabulary for integrating secure practices into the software lifecycle rather than treating them as a final approval event. [[S3]] The same principle applies here: shared controls should be built into delivery, review, and release work instead of recreated for every pilot.",
              "subheading": "Delivery capability"
            },
            {
              "type": "paragraph",
              "text": "The portfolio should remain proportionate. A five-person product team and a public agency do not need identical governance structures. They do need explicit ownership and evidence appropriate to their consequences.",
              "subheading": "Delivery capability"
            }
          ]
        },
        {
          "heading": "Start with decision-ready opportunities",
          "anchor": "start-with-decision-ready-opportunities",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Demand shaping means selecting opportunities based on organizational value and feasibility rather than accepting every attractive idea.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A useful intake asks:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "What decision or workflow is being improved?",
                "Who owns the outcome?",
                "What cannot fail?",
                "Which data and systems are required?",
                "Can representative cases be obtained?",
                "Is a qualified reviewer available?",
                "What is the fallback if AI is unavailable or wrong?",
                "Which shared capability will the work reuse or improve?",
                "What next decision follows the pilot?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Prioritize opportunities where risk and budget already meet: a production release, customer review, failing workflow, modernization boundary, procurement block, or measurable operational backlog.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Use paid pilots as capability-building contracts",
          "anchor": "use-paid-pilots-as-capability-building-contracts",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A paid pilot should specify:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Named business problem",
                "Named decision owner",
                "Bounded users and system path",
                "Time period",
                "Data and access assumptions",
                "Representative cases",
                "Human-review responsibilities",
                "Success, stop, and unknown criteria",
                "Named outputs",
                "Client-owned artifacts",
                "Conversion or next-decision criteria",
                "Explicit exclusions"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The pilot should leave reusable assets even if the system does not proceed. A well-run stop decision can still produce value by preventing a larger failed implementation and clarifying what must change.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Internal capability is not the same as doing everything internally",
          "anchor": "internal-capability-is-not-the-same-as-doing-everything-internally",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Organizations often confuse independence with refusing external expertise. A more useful goal is retained organizational capability.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "External specialists can accelerate discovery, architecture, evaluation, and evidence. The client should still retain:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Authoritative business decisions",
                "System and data ownership",
                "Approved artifacts",
                "Reproducible test cases",
                "Configuration and dependency knowledge",
                "Vendor contracts",
                "Operational runbooks",
                "A clear exit and handoff path"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The anti-lock-in question is not \"Did we use a consultant or vendor?\" It is \"Can we understand, operate, change, and replace this capability without losing the decisions and evidence we paid to create?\"",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Reinvest the AI dividend",
          "anchor": "reinvest-the-ai-dividend",
          "blocks": [
            {
              "type": "paragraph",
              "text": "When AI reduces effort, organizations can extract only the short-term savings or reinvest part of the gain in durability.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Useful reinvestment includes:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Better source data and access controls",
                "Representative evaluation datasets",
                "Human reviewer training and calibration",
                "Modernized integration boundaries",
                "Documentation and system inventories",
                "Operational monitoring",
                "Accessibility improvements",
                "Staff time for learning and process redesign",
                "Secure, reusable delivery components"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The point is not to protect every existing role or process. It is to avoid using efficiency gains in a way that leaves the organization more dependent, less knowledgeable, or less able to manage the next change.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Establish decision rights",
          "anchor": "establish-decision-rights",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A capability program needs clear authority:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Who can sponsor a use case?",
                "Who owns the business outcome?",
                "Who approves data use?",
                "Who defines human-review policy?",
                "Who accepts residual risk?",
                "Who approves production release?",
                "Who can stop or roll back the system?",
                "Who maintains evidence?",
                "Who owns vendor and model changes?"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "Do not place every decision in one central committee. Use federated ownership with common minimum evidence. A central architecture or AI office can define standards and review higher-risk work while product and operational owners retain domain responsibility.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "NIST's AI RMF frames risk management through Govern, Map, Measure, and Manage. [[S1]] Its Playbook offers suggested actions rather than a universal checklist. [[S2]] The same principle applies to an internal operating model: establish shared outcomes and evidence, then scale the process to the actual use case.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Use explicit time horizons",
          "anchor": "use-explicit-time-horizons",
          "blocks": [
            {
              "type": "paragraph",
              "text": "Capability decisions have different time horizons:",
              "subheading": null
            },
            {
              "type": "unordered-list",
              "items": [
                "Now: unblock a release, procurement review, or critical workflow",
                "Next quarter: create reusable evaluation, data, or integration patterns",
                "Next year: reduce platform fragmentation and key-person dependency",
                "Longer term: improve organizational learning, vendor independence, and modernization capacity"
              ],
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "A portfolio should not sacrifice all current value for a theoretical future platform. It should also avoid funding only urgent prototypes that create no durable foundation.",
              "subheading": null
            }
          ]
        },
        {
          "heading": "Measure a balanced system",
          "anchor": "measure-a-balanced-system",
          "blocks": [
            {
              "type": "paragraph",
              "text": "A narrow financial metric can miss capability value. Use four perspectives.",
              "subheading": null
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Buyer and user value",
              "anchor": "buyer-and-user-value"
            },
            {
              "type": "unordered-list",
              "items": [
                "Decision time",
                "Quality and usefulness",
                "Search and source success",
                "Human-review burden",
                "Accessibility",
                "Independent task completion"
              ],
              "subheading": "Buyer and user value"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Internal process",
              "anchor": "internal-process"
            },
            {
              "type": "unordered-list",
              "items": [
                "Release evidence coverage",
                "Failed or blocked cases",
                "Time to reproduce a result",
                "Change-review quality",
                "Incident and rollback performance",
                "Data and access defects"
              ],
              "subheading": "Internal process"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Organizational capacity",
              "anchor": "organizational-capacity"
            },
            {
              "type": "unordered-list",
              "items": [
                "Reusable components and datasets",
                "Documented ownership",
                "Trained reviewers",
                "Reduced key-person dependency",
                "Handoff quality",
                "Ability to replace a vendor or model"
              ],
              "subheading": "Organizational capacity"
            },
            {
              "type": "subheading",
              "level": 3,
              "text": "Financial and decision value",
              "anchor": "financial-and-decision-value"
            },
            {
              "type": "unordered-list",
              "items": [
                "Cost per useful outcome",
                "Avoided rework",
                "Procurement or release readiness",
                "Cost visibility",
                "Value of stopped or narrowed work",
                "Time from trigger to defensible decision"
              ],
              "subheading": "Financial and decision value"
            },
            {
              "type": "paragraph",
              "text": "Do not invent precision where the data is immature. Start with operational signals and improve measurement as the capability becomes real.",
              "subheading": "Financial and decision value"
            }
          ]
        },
        {
          "heading": "Practical operating checklist",
          "anchor": "practical-operating-checklist",
          "blocks": [
            {
              "type": "unordered-list",
              "items": [
                "[ ] Create an intake that names the decision, owner, consequence, data, and next step.",
                "[ ] Distinguish PoCs from paid pilots, readiness, implementation, and operations.",
                "[ ] Define a small shared architecture, data, evaluation, governance, and delivery capability portfolio.",
                "[ ] Require each pilot to reuse or improve at least one shared capability.",
                "[ ] Version datasets, prompts, models, tools, and evidence.",
                "[ ] Name human authority, risk acceptance, stop, and rollback roles.",
                "[ ] Keep client-owned artifacts and exportable handoff materials.",
                "[ ] Reinvest part of efficiency gains in data quality, evaluation, people, and system durability.",
                "[ ] Use explicit near-, medium-, and long-term horizons.",
                "[ ] Measure buyer value, internal process, organizational capacity, and financial decision value.",
                "[ ] Review the portfolio regularly and stop work that cannot reach a defensible next decision.",
                "[ ] Keep external specialists complementary to, not substitutes for, retained client capability."
              ],
              "subheading": null
            }
          ]
        },
        {
          "heading": "Decision connection",
          "anchor": "decision-connection",
          "blocks": [
            {
              "type": "paragraph",
              "text": "The question is not how many AI pilots an organization can launch. It is whether each funded effort improves the organization's ability to make the next decision with less ambiguity, better evidence, and more control.",
              "subheading": null
            },
            {
              "type": "paragraph",
              "text": "The Long-Term Capability Framework provides a five-principle operating model. The Architecture & Reliability Office supports recurring architecture, evidence, and decision review after enough system context exists.",
              "subheading": null
            }
          ]
        }
      ],
      "wordCount": 1470,
      "readingTimeMinutes": 7,
      "limitations": "This operating model must be adapted to the organization's size, risk, procurement model, and technical environment. It is not a universal maturity certification.",
      "sources": [
        {
          "id": "S1",
          "name": "Artificial Intelligence Risk Management Framework 1.0",
          "publisher": "NIST",
          "url": "https://www.nist.gov/itl/ai-risk-management-framework",
          "accessed": "2026-07-22",
          "note": "Government framework"
        },
        {
          "id": "S2",
          "name": "NIST AI RMF Playbook",
          "publisher": "NIST",
          "url": "https://airc.nist.gov/airmf-resources/playbook/",
          "accessed": "2026-07-22",
          "note": "Government implementation resource"
        },
        {
          "id": "S3",
          "name": "Secure Software Development Framework, NIST SP 800-218",
          "publisher": "NIST",
          "url": "https://csrc.nist.gov/pubs/sp/800/218/final",
          "accessed": "2026-07-22",
          "note": "Government framework"
        }
      ],
      "sourceMarkers": [
        "S3",
        "S1",
        "S2"
      ],
      "implementationChecklist": [
        "[ ] Create an intake that names the decision, owner, consequence, data, and next step.",
        "[ ] Distinguish PoCs from paid pilots, readiness, implementation, and operations.",
        "[ ] Define a small shared architecture, data, evaluation, governance, and delivery capability portfolio.",
        "[ ] Require each pilot to reuse or improve at least one shared capability.",
        "[ ] Version datasets, prompts, models, tools, and evidence.",
        "[ ] Name human authority, risk acceptance, stop, and rollback roles.",
        "[ ] Keep client-owned artifacts and exportable handoff materials.",
        "[ ] Reinvest part of efficiency gains in data quality, evaluation, people, and system durability.",
        "[ ] Use explicit near-, medium-, and long-term horizons.",
        "[ ] Measure buyer value, internal process, organizational capacity, and financial decision value.",
        "[ ] Review the portfolio regularly and stop work that cannot reach a defensible next decision.",
        "[ ] Keep external specialists complementary to, not substitutes for, retained client capability."
      ],
      "relatedCapabilities": [
        "ai-production-evaluation",
        "application-data-integration",
        "technical-evidence-procurement"
      ],
      "relatedServices": [
        "architecture-reliability-office",
        "ai-production-readiness"
      ],
      "relatedEvidence": [
        "governed-knowledge-system",
        "public-rd-governed-handoff"
      ],
      "relatedResources": [
        "long-term-capability-review",
        "paid-pilot-charter-template"
      ],
      "seo": {
        "description": "AI programs compound when they reuse architecture, data boundaries, evaluation, governance, ownership, and learning. Isolated pilots usually repeat the same discovery and risk work.",
        "socialTitle": "Build reusable capability, not a collection of AI pilots"
      },
      "status": "published",
      "contentOwner": "Mike Kappel"
    }
  ]
}
