Guide / Delivery

Six-week production AI delivery plan

Six weeks is enough to ship one bounded AI workflow when each week produces working evidence and the customer owns the environment, decisions, and handoff.

September 3, 2026

Six-Week Production AI Delivery Plan: Scope, Build, Rollout, and Handoff

A six-week delivery window can produce a real system. It cannot rescue an unbounded mandate such as "add AI to support," replace missing executive decisions, or compress every security and procurement process into the final week.

The plan works when one senior engineer owns one measurable workflow inside the customer's environment. Each week must end with code, data, evaluation evidence, or an operating decision that the customer can inspect.

What must be true before week one

Choose one workflow with a named business owner, a clear user, available data, a production path, and a measurable outcome. Examples include drafting a support resolution from approved knowledge, extracting fields from one document class, or preparing an analyst brief from governed sources.

Write the boundary in one sentence:

For this user, the system receives these inputs, performs these allowed actions, produces this output, and stops here.

Record a baseline from the current process. Measure the outcome, not only model quality. Useful measures include completion rate, handling time, correction rate, escalation rate, cost per completed case, and user adoption.

AWS guidance for a generative AI proof of concept recommends validating business value, data readiness, technical feasibility, and delivery risk before a go decision. Use those as entry conditions. If access to required data, a process owner, or a deployable environment is missing, the six-week clock has not started.

The operating team

A compact delivery team needs explicit decision rights:

RoleOwnsWeekly evidence
Business ownerOutcome, policy, user adoptionMetric and acceptance decisions
Forward deployed engineerEnd-to-end technical deliveryWorking code and gate evidence
Domain reviewerCorrectness and edge casesLabeled eval examples
Platform ownerIdentity, data, deploymentApproved integrations and controls
Risk ownerPrivacy, security, complianceWritten constraints and sign-off
Operations ownerMonitoring and responseAlerts, runbook, on-call acceptance

One person may hold several roles. No role may be implied. If a gate needs approval, name who can approve it and when they will be available.

Week 1: Lock scope, baseline, and architecture

Map the current workflow from trigger to final disposition. Observe real users and collect representative cases. Identify inputs, systems of record, decisions, exceptions, downstream actions, and the human who remains accountable.

Create the first evaluation set from actual workflow examples. Include normal cases, high-value cases, ambiguous inputs, malformed inputs, sensitive data, unavailable tools, and conditions where the system must refuse or escalate. Domain reviewers should label acceptable results before the team optimizes against them.

Choose the smallest architecture that can prove the workflow. Prompting may be enough for a closed transformation. Retrieval is appropriate when answers need approved changing knowledge. Tool-using agents add value when the workflow requires bounded actions, but they also add permissions and failure modes.

Deliver by Friday:

  • Signed workflow boundary and out-of-scope list.
  • Baseline business and system metrics.
  • Data inventory and access owners.
  • Initial evaluation set and scoring rubric.
  • Architecture decision record.
  • Threat and failure-mode review.
  • Six-week release and stop criteria.

Gate 1: proceed only if the data is usable, the integration path is real, and the evaluation set can distinguish an improvement from a demo.

Week 2: Ship the thinnest vertical slice

Build one complete path from real input to reviewable output in the target repository. Connect one identity path, one source of truth, one model route, and one result surface. Use a non-production environment, but preserve the production boundaries.

Instrument the slice now. Capture request IDs, model and prompt versions, retrieval references, tool calls, latency, token usage, errors, and reviewer outcomes. Logs must avoid exposing protected content and must support reconstruction of a failed case.

Run the evaluation set in CI or another repeatable harness. Separate task quality from system behavior. A high-quality answer that bypasses authorization, times out under normal load, or costs more than the workflow can support is not a passing result.

Deliver by Friday:

  • Deployed vertical slice in a controlled environment.
  • Automated evaluation harness with versioned cases.
  • Trace and cost instrumentation.
  • Demonstration using representative, not hand-picked, cases.
  • Updated architecture and risk log.

Gate 2: the complete path works for representative cases and every result can be traced to its code, configuration, model, and evidence.

Week 3: Integrate the real workflow

Replace mock boundaries with production-shaped integrations. Implement identity propagation, least-privilege authorization, input validation, data retention, rate limits, timeouts, retries, idempotency, and human approval before consequential actions.

Design failure paths deliberately. When retrieval returns weak evidence, the system should ask for clarification or route to review. When a tool is unavailable, it should preserve state and explain what did not happen. When a downstream action is rejected, the operator should be able to retry safely.

AWS production guidance recommends promoting prompts, model settings, code, evaluation data, and environment configuration as coordinated, versioned assets. Treat a release as a tested bundle. Do not update the prompt, model, and retrieval pipeline independently without recording which combination passed.

Deliver by Friday:

  • Real integration contracts and access controls.
  • Tested error, timeout, retry, and idempotency behavior.
  • Versioned release bundle.
  • End-to-end evaluation results.
  • Draft operator dashboard.

Gate 3: the workflow can fail safely and can be operated without the delivery engineer inspecting raw application state.

Week 4: Harden quality, security, and operations

Expand evaluation from the failures found in weeks two and three. Add adversarial and boundary cases. Test prompt injection, unauthorized tool requests, data leakage, unsupported claims, duplicate actions, partial upstream data, provider outage, and latency under expected concurrency.

Set explicit thresholds for quality, safety, reliability, latency, and cost. Averages can hide severe failures, so inspect tail latency and subgroup performance. Decide which failures block release and which create a human-review route.

Build production controls: dashboards, alerts, budget limits, kill switch, rollback, degraded mode, model and prompt pinning, and audit events. Run an incident exercise in which the primary model or a critical data source becomes unavailable.

Google Cloud's generative AI MLOps blueprint treats deployment, evaluation, monitoring, and governance as one lifecycle. That is the practical standard for this week: the system is not ready until the team can observe and control it.

Deliver by Friday:

  • Expanded eval suite with regression thresholds.
  • Load, failure, and security test evidence.
  • Monitoring and alert routing.
  • Rollback and degraded-mode proof.
  • Draft production runbook.

Gate 4: every release blocker has passing evidence, and operations has demonstrated containment and rollback.

Week 5: Canary with real users

Release to a small, named user group or a low-risk slice of traffic. Keep the prior workflow available. Train users on the system's boundary, the review responsibility, and the escalation path.

Compare results with the week-one baseline. Track both technical and business measures. Watch correction rate, abstention, escalations, adoption, latency, cost, and any harm indicator. Review failures daily with the domain owner and turn confirmed cases into permanent evaluations.

Use predetermined expand, hold, and rollback thresholds. Avoid expanding because feedback "feels positive." Expand only when the recorded evidence clears the gate.

Deliver by Friday:

  • Canary release and cohort record.
  • Live metrics compared with baseline.
  • User feedback and correction analysis.
  • Updated evaluation set.
  • Release, hold, or rollback decision.

Gate 5: the canary meets the outcome and safety thresholds for a defined observation period, and the business and risk owners sign the expansion decision.

Week 6: Release and transfer ownership

Promote the exact tested bundle. Expand in stages, verify metrics at each stage, and keep rollback available. Record the deployed code revision, configuration, prompts, model route, evaluation result, approvals, and change window.

Finish the runbook around decisions an operator must make:

  • How to confirm the system is healthy.
  • What each alert means and who owns it.
  • How to disable an action or switch to degraded mode.
  • How to roll back code, prompts, models, and indexes.
  • How to replay or repair a failed case safely.
  • How to update the evaluation set.
  • How to approve a future release.
  • What evidence is retained for audit.

Transfer repository, deployment, dashboards, alert routes, vendor accounts, cost controls, evaluation assets, architecture decisions, and known limitations to the customer's named owners. Pair on a release and an incident drill. The handoff succeeds when the customer performs the procedure while the embedded engineer observes.

Microsoft's AI adoption planning guidance emphasizes aligning strategy, organizational readiness, governance, and operations. The transfer should leave those responsibilities visible inside the client team.

Deliver by Friday:

  • Controlled production release.
  • Signed acceptance against the original measures.
  • Complete runbook and architecture record.
  • Owner matrix and on-call routes.
  • Client-run deployment and rollback exercise.
  • Prioritized post-release backlog.

Gate 6: the system is operable by the customer's team, production evidence is retained, and rollback has been demonstrated.

Evidence packet for every Friday

Keep the weekly review short and inspectable. Bring:

  1. The current production candidate and commit.
  2. The evaluation result with failures, not only the headline score.
  3. Business metric movement against baseline.
  4. Reliability, latency, and cost measures.
  5. New risks and decisions with owners.
  6. The next gate and its exact pass condition.

A slide deck can summarize this packet. It cannot replace working evidence.

When to stop, narrow, or extend

Stop when the workflow does not create measurable value, required data cannot be used lawfully or reliably, a critical integration cannot be controlled, safety thresholds cannot be met, or no client owner accepts operations.

Narrow scope when one exception path causes most of the risk, a smaller user cohort can validate the outcome, or a human approval step can contain uncertainty.

Extend the schedule when a mandatory external process has real lead time, such as security review, procurement, identity integration, regulated validation, or collection of a representative evaluation set. Preserve the gates rather than declaring production on the calendar.

The six-week promise is a delivery constraint, not evidence by itself. The outcome is one bounded system with working code, repeatable evaluations, observable behavior, an operational runbook, and owners who can deploy, stop, and improve it.

Sources

Have one AI workflow that needs to reach production?

FDE embeds a senior engineer in your team and repository to scope, build, harden, roll out, and hand off a bounded production workflow.

Book a 15-minute scoping call