AI readiness for small-business managers — Part 3 of 5.

Every AI deployment plan contains the sentence “a human will review the output.” Parts 1 and 2 gave
you a boundary and an evaluation; this part tests that sentence, because it is where most plans
quietly stop being true. There is a difference between a named reviewer and a real intervention
system, and the difference is five requirements you can check in an afternoon.

Step 1 — test your reviewer against the five requirements

From the training: a reviewer who can actually intervene needs

  1. access to the source — they can see the document the AI read, not just the AI’s answer;
  2. authority and time to reject the output — saying no is part of their job, and the queue
    allows it;
  3. a way to record corrections — every fix is written down somewhere that survives the day;
  4. an escalation route — a known person to call when the output is worrying rather than wrong;
  5. protection from automation bias and impossible workloads — because a reviewer facing 400
    items a day approves; they do not review.

Check all five for your service. Any requirement you cannot evidence goes on the Part 1 list as
Unverified, with an owner and a date.

In August 2026 I hired Gemini and fired it the next day. The company record says it made up facts
about one of my shops. What I remember is the behaviour. It says, okay, I’m going to do this, and
it never does it. It says it cannot find the information in the chat when it’s actually there. I
give it a command and it does something completely different. I asked it for a three-second video
and it refused; I told it I had the people’s agreement, and it generated the previous video again
and charged my tokens for it. At that time the AIs didn’t collaborate. I was literally carrying
instructions: Claude told me what to ask, I transferred the request to Gemini, then I told Claude
what the answer was. That was the oversight, and the only thing I could stop was the whole
arrangement. Since then I have this model of checking Claude’s work with ChatGPT, and eventually my
human eye decides.

— the five requirements as a checklist card, worked-example answers beside them

The fifth requirement deserves the extra minute. Automation bias is the tendency to accept a
confident machine answer, and it grows with workload and with trust — the reviewer who caught
errors in week one waves them through in week six. The checklist treats it as a design problem:
interface, training and sampling, not a memo asking people to stay alert.

Step 2 — find the intervention point

The training’s exercise: take one output of your system that would matter if wrong — a price, a
classification, a summary a decision rests on — and trace its path to the decision. Mark where a
human could still stop it before harm, not after. If the honest answer is “nowhere until the
customer complains”, you have monitoring, not oversight.

— the worked-example trace: document in, output, review point, decision, with the intervention point marked

Step 3 — answer the stop question

The training’s sharpest test comes from the governance session: who can stop the system today?
Not who could convene a meeting about stopping it — who can pause processing this afternoon, on
their own authority. From the RACI’s decision rights, a small business needs four names:

Decision Who holds it
Accepts residual risk The business owner — a name, not a committee
Approves routine changes The implementation lead, inside the documented boundary
Approves material changes (purpose, model, data, automation level) The business owner
Orders an immediate safety rollback or pause The technical owner, without waiting for anyone

In a five-person company the same person may hold two of these. That is fine and normal; what the
governance rules refuse is the reverse — a decision nobody holds. If the answer to the stop
question is unclear, the control is missing, whatever the org chart says.

— the four-decision card with worked-example names filled in

Step 4 — make corrections count

The requirement managers most often skip, from checklist domain 6: human corrections are captured
as monitoring evidence rather than disappearing inside the workflow. When a reviewer fixes an
output and moves on, the organisation learns nothing; when the fix is logged, the correction rate
becomes your earliest warning that something changed. Part 4 builds the thresholds on exactly this
data — a correction log is oversight feeding monitoring, one artefact doing two jobs.

The log needs three columns and no software: what was corrected, on which document type, by whom.

Step 5 — check the workload arithmetic

From the diagnostic: is reviewer workload realistic at peak volume? Multiply expected daily items
by honest minutes-per-review and compare the result with the hours that exist. If review time
exceeds the time the AI saves, the case for the system changes — and it is better to discover that
in arithmetic than in a burned-out reviewer rubber-stamping the queue. Oversight that only works
at demo volume is a pilot constraint, and it belongs in the launch decision in Part 5.

What you have when you finish

A tested reviewer role instead of an assumed one: five evidenced requirements, a marked
intervention point, four named decision holders, a correction log that feeds monitoring, and
workload arithmetic that survives peak volume.

Part 4 takes the operational half of the same machine — what to monitor once the service runs,
the five warning thresholds that trigger investigation, and what has to happen when one fires.


Authorship: HAC — human-directed, AI-assisted. The passage on hiring and firing Gemini, in Step 1, is the author’s own spoken words, transcribed and arranged, not generated. This part is converted from the 90-minute manager training “From AI prototype to production service” (segment 40–55 min), the Day 10 readiness checklist (domain 6), the Day 9 governance RACI, and the readiness diagnostic (oversight section). No new frameworks were created for it.

The AI readiness course: Part 1 — The prototype trap · Part 2 — Evaluation that can fail · Part 3 — Meaningful human oversight · Part 4 — Running it for real · Part 5 — The launch decision.