06

A weekly Founder Note on the infrastructure, decisions and operating models shaping enterprise AI.

When a company says its AI system has a human in the loop, I ask a simple question: which human, at which point, with what information?

Too often, the answer is vague. A person is expected to catch a mistake after the system has already produced a confident response, or to approve a queue of cases without seeing the evidence behind them.

That is not meaningful oversight. It is a handoff without a design.

Enterprise AI needs human judgment where the consequences, uncertainty or context demand it. That judgment has to be built into the workflow from the start.

Decide what the system may actually do

The first decision is about authority. An AI assistant that drafts a customer reply is different from one that sends it. A system that summarizes a claim is different from one that settles it.

For each workflow, define three levels: tasks the system can complete, tasks it can recommend for approval, and decisions reserved for a person. The boundary should reflect the impact of an error, the quality of available evidence and the ability to reverse the outcome.

This turns oversight from a general principle into an operating rule. A team can then measure how often the system stays within that rule and where it needs to change.

Escalation needs a useful reason

Sending every uncertain case to a person can overwhelm a team. Sending none can expose customers to avoidable harm. The answer is to define specific triggers.

A missing source, conflicting policies, an unusual request, a sensitive customer circumstance or a language the system handles poorly may each justify review. A confidence score alone is rarely enough; it must be calibrated against real outcomes and the particular task.

The reviewer should receive the original request, the source material, the proposed action and the reason for escalation. Otherwise the handoff forces a person to reconstruct the case from scratch.

Make corrections part of the system

A reviewer should be able to correct the answer, identify a bad source, change the workflow rule or flag a new exception. Those are different problems and require different fixes.

If the same correction appears repeatedly, the organization should not ask employees to keep fixing individual outputs. It should update the approved knowledge, terminology, evaluation set or escalation logic. Human oversight creates value when it improves future decisions.

Language changes the risk of a handoff

Consider a customer explaining a disputed bank transaction in a regional language. Speech recognition may mishear a name or amount. Machine translation may flatten a distinction between a question and an authorization. A language model may produce a polished answer that hides the uncertainty.

For sensitive interactions, a reviewer needs access to the original audio or text alongside the transcript, translation and source-backed recommendation. Approved terminology and bilingual review help preserve meaning. The customer should not have to switch to English for the organization to take their concern seriously.

The purpose is not to make a person translate every conversation. It is to place review where a language error could change the decision.

Measure the quality of both sides

Automation rate is only one measure. Track the share of cases escalated, the reasons, review time, corrections, customer outcomes and errors found after completion. Look for cases that should have been escalated and were not.

Reviewers also need clear ownership, training and enough time to challenge the system. If approval becomes a reflex, the presence of a human does not make the process safer.

The strongest AI systems make people more effective and use their feedback to become more reliable.

What next

What this means for enterprise leaders

Take one AI workflow and draw its decision boundary. Mark what the system can do, what requires approval and what must remain a human decision.

Then inspect ten real or realistic cases. For each escalation, ask whether the reviewer had the original context, the right evidence and a clear way to correct the underlying problem.

In the next Founder Note, I will examine why multilingual capability must be designed across the entire AI workflow, rather than added as a translation step at the end.

Himanshu SharmaCo-founder, Devnagri AI