Mary Fung
essayOctober 1, 2026

Simulation is a gym for judgment

If AI removes some of the repetitions through which people learn, practice needs a new home.

The usual argument for simulation is safety.

Do not let the new person make their first difficult decision with a real customer, a real patient, a real payment, or a real system outage. Give them a bounded version of the situation first. Let them see the consequences when the cost of being wrong is still low.

That is right. It is also too small.

As AI takes over more of the routine work through which people used to learn, simulation may become part of the curriculum.

Not because the old work was beautiful. Much of it was boring, wasteful, and poorly designed. But it contained repetitions. Repetitions produced mistakes. Mistakes produced corrections. Corrections eventually became the quiet feeling that something was off before a person could quite explain why.

That feeling is judgment.

We are getting very good at giving people the answer. We are much less clear on how they will get the repetitions that once taught them to question it.

The first draft was doing two jobs

Take a new analyst asked to prepare a recommendation.

Historically, producing the first version took time. They had to find the relevant information, decide what mattered, make assumptions visible, and get something wrong in a way a more experienced person could correct. The final document was one output. The developing judgment was another.

An AI system can now make the first output arrive much faster. That can be an enormous improvement. It may expose a junior person to better examples and let them work on more difficult questions sooner.

But it does not automatically do the second job.

The person may receive a fluent recommendation before they have had to form one. They may learn what a finished answer looks like without learning what evidence it rests on, where it can fail, or what competing interpretation it excluded.

The risk is not simply that people forget how to make a document from scratch. It is that they become reviewers of work they have not learned to judge.

There is no prize for preserving pointless manual labour. The question is narrower: which experiences were quietly building a capability that the new workflow still needs?

A gym is not the same as a test

Most organisations already know how to test whether a system behaves as expected. A team can assemble cases, run the system against them, record errors, and improve the system.

That is a useful test harness. It is not yet a learning environment for a person.

A gym has repetitions designed around the person who is training. It makes difficulty visible. It gives feedback before the person has forgotten the decision they made. It lets them try a different approach. It does not reward getting an answer right by luck and moving on.

An AI-era simulation should do the same.

Give someone a realistic but bounded scenario. Ask them to form a view before the assistant offers one. Let the assistant surface missing evidence, opposing interpretations, or similar cases. Show what follows from a choice. Let the person explain why they would still make it. Then change one condition and see whether their reasoning transfers.

The point is not to catch them out. It is to make the hidden work of judgment available for practice.

That is different from putting a chatbot beside a training module and calling the experience personalised.

Put AI in the right part of the loop

The same AI can either weaken or develop judgment. The difference is where it enters.

If it appears before a person has considered the problem, it can become a substitute for attention. The answer looks reassuringly complete. The person learns to recognise its tone before they learn to interrogate its reasoning.

If it appears after a person has formed a hypothesis, it can become an unusually patient sparring partner.

It can ask what assumption is carrying the conclusion. It can produce the strongest case against the recommendation. It can show how a previous case differed in the one detail that now matters. It can introduce a constraint that makes an apparently sensible answer fail.

None of that means the system has judgment. It means the system can create more occasions for a person to exercise theirs.

This is what productive friction looks like in practice. Not more bureaucracy. Not a ritual human approval at the end. A moment early enough in the work for a person to notice, revise, and learn.

The simulation needs disagreement

One risk of a well-designed AI assistant is that it becomes too agreeable.

It can produce a coherent answer to almost any prompt. It can make a person feel that they have explored a question because the conversation was long. But more thinking inside the same frame does not necessarily reveal that the frame is wrong.

The useful simulation includes another perspective.

Sometimes that perspective is a colleague who knows the work differently. Sometimes it is a recorded failure case. Sometimes it is an AI instructed to argue from the constraint the first answer ignored. The mechanism matters less than the result: the person should encounter a plausible reason to revise their view.

Teams need this too. If everyone completes private work with their own assistant, they may become faster while building less shared understanding. A simulation can make the assumptions visible before they harden into a plan. It gives a team somewhere to disagree while the decision is still cheap.

That is why a gym needs other people in it. You can lift alone. You cannot reliably see your own form.

What to measure instead of completion

The easiest measure for a training system is completion. Did the person finish the module? Did they use the tool? Did they choose the expected answer?

Those measures are not useless. They are just too weak.

The better questions are:

  1. Can the person explain the evidence behind the recommendation?
  2. Can they name the condition that would change their mind?
  3. Can they spot the relevant difference when the familiar pattern no longer fits?
  4. Can they use the assistant to challenge a view, rather than only to generate one?
  5. Can they carry the reasoning into a new case?

This is harder to measure than tool usage. It is also closer to the capability the organisation will need when the system is wrong, unavailable, or faced with a situation it has never seen.

The old apprenticeship happened accidentally. Useful work and human development were entangled, so no one had to design the curriculum explicitly.

AI is separating them.

That is not necessarily a loss. It may be an opportunity to build a better curriculum: less drudgery, more deliberate practice, faster feedback, and more exposure to the cases that actually form judgment.

But the curriculum will not appear automatically because the assistant is clever.

Someone has to decide that the person matters as much as the output.

← back to the field