AI readiness for small-business managers — Part 4 of 5.

Parts 1–3 built the service’s paperwork: boundary, evaluation, oversight. This part is about the
day after launch, because a production AI service is not monitored only for uptime — it must be
watched for whether its answers stay supported, whether humans are correcting more of its work,
and whether a supplier update quietly made it worse. None of that requires a platform. It requires
five monitoring habits, five thresholds, and named owners.

Step 1 — put names on the roles

From the RACI, six roles carry a production service. In a small business several land on the same
person; write the names anyway, because the gap you are closing is the decision nobody holds.

  • Business owner — owns the outcome, funds it, accepts residual risk.
  • Implementation lead — coordinates requirements, evaluation and controls.
  • Technical owner — keeps it running, operates monitoring, can roll it back immediately.
  • Data/privacy lead — owns lawful data handling and commands personal-data incidents.
  • Compliance/risk — interprets obligations and challenges the controls.
  • Frontline reviewer — reviews output and supplies the correction evidence.

The supplier is deliberately not on the list of owners. It supports, notifies and provides
evidence, but its assurance is an input to your decision, not a substitute for it — the supplier
cannot accept your risk on your behalf.

— the six-role card with worked-example names, several roles sharing one person

Step 2 — watch five things

The Day 8 framework monitors five families. Each is a question a manager can ask weekly:

  1. Output quality — are errors appearing in critical fields, and are they severity-graded
    rather than averaged?
  2. Human corrections — is the correction rate rising, and on which document types? (This is
    the log from Part 3 doing its second job.)
  3. Coverage — are new layouts, languages or out-of-boundary uses arriving that the evaluation
    never saw?
  4. Performance and cost — response times, queue length, cost per item against budget.
  5. Change and drift — is every output traceable to the model, prompt and integration versions
    that produced it, and did the regression set rerun after the last change?
— the five families as a one-page monitoring sheet

Step 3 — set the five warning thresholds

The framework’s provisional triggers, to calibrate against your own baseline before production:

  1. Critical-field error above 2% in sampled output — or a single error that could carry a
    legal, financial, safety or personal-data consequence — pause automated processing and
    investigate.
  2. Correction drift: seven-day correction rate above 10%, or a sustained rise of 50% over the
    approved baseline.
  3. Unsupported or missing output above 5% for any document type, language or user group,
    even while the overall average looks fine.
  4. Latency or cost regression: p95 latency 30% over target, or cost per item 25% over budget,
    for three days.
  5. Change regression: the fixed evaluation set falls more than five percentage points, or a
    previously closed high-severity failure returns — roll the change back.

The numbers are starting points; the structure is not negotiable. A threshold without a named
investigation owner and a response time is a dashboard, and dashboards have never paused a system.

— the threshold sheet with owner and response-time columns filled in

Step 4 — control change

From checklist domains 4 and 8: version the model, prompts, retrieval sources and integrations;
require evaluation, approval, release notes and a tested rollback route for every material change;
and treat supplier changes as changes — reviewed, not merely received. The decision rights from
Part 3 apply here: routine changes are the implementation lead’s, material changes are the
business owner’s, and the technical owner can roll anything back on the spot.

Step 5 — prepare for the bad day

My bad day was a password. Claude told me, at some point, that a WordPress admin password was
compromised: it was sitting in a handover file that had been pushed to the repository. I don’t
quite remember exactly what and why. We had been working locally, the password was local, and I
think it was a very easy one. From “this is wrong” to “this is fixed” it lasted a while, because I
didn’t take the decision on the spot. Claude always tells you there are compromised passwords, and
it still makes the mistake of exposing them in chats. So whenever it tells me about another one,
I’m thinking: is it an important one? Is it really as Claude says? If I have to, I change it, but
not all the time. If a hacker can get into my laptop, he can read all my chats and the hard drive
anyway, so I don’t know if that’s a big problem. What I did change is the way we keep them: in a
folder on the hard drive, where Claude or ChatGPT can take them without putting them in a chat. I
don’t remember the day the machine check went in. It’s down to Claude to fix this at some point.

From checklist domain 10, five requirements, all boring until needed:

  • Staff know how to report an error, suspected harm or data incident.
  • Operational incidents have a named commander (technical owner); personal-data incidents have

their own (data/privacy lead).

  • Pause, restrict, roll back and suspend can be executed without waiting for the supplier.
  • Recovery has criteria: what evidence is required before normal operation resumes.
  • The manual fallback can sustain the business for an agreed period — tested, not assumed.

Every threshold breach produces a short record: what fired, which versions were live, what was
affected, the containment action, the named owner, and the decision — continue, restrict, fall
back, roll back or suspend.

— the minimum investigation record as a fill-in card

What you have when you finish

Six named roles, a five-line weekly monitoring habit, five owned thresholds, change control with a
rollback route, and an incident plan that works while the supplier’s ticket queue does not. In the
worked document-agent case this entire layer was Unverified — no baseline, no alert ownership, no
tested fallback — and that finding, more than any accuracy number, is what kept the case at pilot.

Part 5 puts everything on one page and makes the decision the course exists for: Go, Conditional
go, Pilot only, or No-go.


Authorship: HAC — human-directed, AI-assisted. The opening of Step 5, on the bad day, is the author’s own spoken words, transcribed and arranged, not generated. This part is converted from the 90-minute manager training “From AI prototype to production service” (segments 55–80 min), the Day 10 readiness checklist (domains 4 and 7–10), the Day 8 monitoring and drift framework, and the Day 9 governance RACI. No new frameworks were created for it.

The AI readiness course: Part 1 — The prototype trap · Part 2 — Evaluation that can fail · Part 3 — Meaningful human oversight · Part 4 — Running it for real · Part 5 — The launch decision.