Dr-Business Start the fit check

Stop Assigning AI the Judgment Calls It Can’t Make

To decide who should perform a task, find its last action: the step that makes the work real in a system, in front of a customer, or in a decision someone must defend. Then ask how much judgment that action requires and how easily it can be undone.

Those two properties allocate the work more reliably than capability alone. A tool may produce an excellent output and still be the wrong party to perform the consequential step.

Capability is only one axis

Most teams allocate by capability. Something gets tried, the output looks good, and the task moves across. The reasoning is sound as far as it goes: a tool that produces a usable draft has demonstrated it can produce a usable draft.

What that reasoning omits is consequence. A model may draft a routine reminder and a formal notice with equal fluency. The two outputs may read equally well, while one can be corrected quietly and the other can end a tenancy.

Capability and consequence vary independently, so sorting by one tells you very little about the other. That is why teams get surprised: the failure arrives in a task the tool handled beautifully, in a place where handling it beautifully was beside the point.

The Last Action

Take any task and identify its last action: the step that turns an output into an operational fact. It might create a database record, send a message, commit the business to a decision or change a legal relationship.

Then rate that action on two axes. How much judgment does it require? And if it is wrong, how cheaply, quietly and quickly can it be undone?

The combination places the task in one of four modes. That is The Last Action.

These questions are usually easier to answer than an abstract question about what AI can do, because the team can inspect where the action lands and what reversing it would require.

What follows uses one setting throughout: a property management company handling residential units. All four modes appear inside that one business, which is the point — they are properties of tasks, not of companies.

The four modes

Runs unattended

Maintenance requests arrive by email and by message, all day, in whatever wording the tenant chose. Something reads each one and files it against the right unit with a category and a date.

The last action here is a database row. If the category is wrong, someone corrects it when they open the ticket, and no one outside the company ever saw the error. The action is reversible, quiet, and cheap, and the judgment involved is low.

This mode is often over-supervised. When the last action is a reversible internal record, continuous human review may cost more than the errors it prevents.

Drafts, a person sends

The tenant needs a reply: the plumber comes Thursday between nine and one, here is what to expect. Drafting this well takes context about the unit, the history and the tone the company uses, and a model with that context drafts it very well indeed.

The last action is a message to a customer. Reversibility drops sharply — a correction can follow, but the first message has already been read, and the reader forms an impression the correction only partly retrieves.

So the draft is produced automatically and a person presses send. This is the mode where drafting quality is highest and delegation of the final step is still wrong, and it is the clearest evidence that capability alone is insufficient. The tool is excellent at the work. The person is there for the consequence.

A person decides, AI supports

A repair is expensive enough to trigger a dispute, and the tenancy agreement is ambiguous about whether wear of this kind sits with the landlord or the tenant. Someone has to decide.

Here the judgment is genuinely high — it involves the agreement, the relationship, what was decided in a similar case last year, and what the landlord will accept. A model can assemble all of that: pull the relevant clause, surface the earlier case, lay out the two readings and what each implies. That is real work and it saves real time.

The decision stays with a person, and the reason has little to do with the quality of the analysis. It is that the person carries the relationship afterwards and will have to defend the call to two parties who both have standing to argue with it.

A person only

A notice to vacate. A termination. A statement to a landlord that their unit will be off the market for six weeks.

The last action is irreversible in the plain sense — legally, commercially, or in the relationship — and the judgment is high. In this example, even the first draft stays with a person. Fluent wording can anchor the reviewer before they have fully considered the legal, commercial and relationship consequences.

That does not make every important document person-only. This task belongs here because both the judgment and the final action are difficult to reverse.

This mode should be small. If it holds most of a team’s work, the classification is being done by anxiety in place of the two questions.

When reversibility is unclear

The honest complication is that reversibility often depends on context that is easy to miss in advance.

A routine reminder sent to the wrong tenant is embarrassing and recoverable. The same reminder sent to a tenant already in dispute is a document that reappears in a legal letter three months later. The message is identical; the reversibility differs, and the difference lives in context the sender may lack at the moment of sending.

Where the consequence varies by recipient, split the task into separate cases if the relevant context can be detected reliably. If it cannot, classify the set according to the higher plausible consequence.

Where a task remains genuinely borderline, place it one mode closer to the human for a limited period and review what actually happens. The cost of over-supervising a reversible task is measured in minutes. The cost of the reverse is measured in relationships.

These are modes, and not a ladder

A tempting reading of the four is that they are stages, and that a maturing team walks tasks upward until most things run unattended.

That reading is wrong and it is worth naming, because it produces exactly the wrong behaviour. The modes are properties of tasks. A notice to vacate belongs in the fourth mode permanently, in a company using AI heavily and in one using it barely, because the last action ends a tenancy in both. Ticket categorisation belongs in the first mode on day one.

A team improves by classifying accurately, and a correct classification often moves a task toward the human. Both directions are progress.

Sort five tasks this week

Take five tasks a model already touches in your business and identify the last action in each. Then ask: how much judgment does that action require, and how easily could it be undone if it were wrong?

Write the resulting mode beside each task. Expect at least one task to move in each direction: a reversible internal action that is being over-supervised, and an external action that should return to a person even though the drafts have been good.

The second of those is the one worth finding first. It is the task that has been working fine every day and will produce the surprise, because the output was always good and the question was always about something else.

Before you bolt on another tool, it is worth knowing whether your business runs on systems or on you. I put together a free 2-minute assessment that gives you a straight read on exactly that, and the first thing to fix. Take the free assessment.

WORK WITH US

Ready to make your AI actually reliable?

Book a diagnosis and we will map the highest-leverage fixes for your business.

Book a diagnosis
NEWSLETTER

Sharper signal. Smarter decisions.

Join our newsletter for our best thinking on AI and systems, delivered straight to your inbox - no noise.

Subscription Form
No spam. Unsubscribe anytime.

Related posts

Leave the first comment