Executives should not be comfortable with unlimited AI spend and no ROI.
That is not strategy. That is a tab left open.
But they also should not demand immediate margin from every token the way they would from a mature operating process. That creates a different failure mode: the organization becomes so afraid of waste that it never learns where the value is.
The problem is not that AI spend is hard to justify.
The problem is that companies keep mixing two different kinds of spend and then pretending one ROI question can govern both.
Some AI spend is learning spend.
Some AI spend is production spend.
They need different rules.
Learning spend should reduce uncertainty
Learning spend is the cost of finding out.
It is closer to R&D than operations. The point is not immediate margin. The point is reducing uncertainty faster than competitors, vendors, committees, or quarterly planning cycles can.
The organization is paying to learn:
- Which workflows are actually AI-suitable?
- Which teams can use AI well?
- What context is missing?
- Where does quality break?
- What needs human review?
- Which model or tool is good enough for this task?
- Which work becomes easier but not valuable?
- Which old process should stop if the AI-assisted version works?
Those are real business questions.
They are just not always answered by a clean financial return in the first week.
This is where many executives get trapped. They know enough to dislike vague innovation spend. They are right to dislike it. But the alternative cannot be demanding that every experiment prove full ROI before the organization has learned what is worth measuring.
That is how companies end up funding safe theater.
A low-risk demo can look more justifiable than a messy workflow experiment. A prompt workshop can look cleaner than changing how review happens. A dashboard of usage can look more mature than admitting that the organization still does not know where AI should be allowed to act.
Learning spend should be allowed.
It should also be capped, time-boxed, and judged by the decisions it produces.
The right question is not:
Did this experiment already return margin?
The right question is:
What did this spend teach us that changes what we do next?
If the answer is nothing, stop.
Production spend needs unit economics
Production spend is different.
Once AI is inside a recurring workflow, the organization should stop treating the spend as exploration and start treating it as a cost of production.
At that point, token spend belongs next to the unit of work it supports:
- cost per resolved support case
- cost per reviewed contract
- cost per generated report
- cost per qualified lead
- cost per research memo
- cost per code migration
- cost per training module
- cost per customer interaction
- cost per decision cycle
This is where discipline gets much easier.
The question becomes:
Is the AI-assisted workflow cheaper, faster, better, safer, or newly possible after AI costs are included?
Not after pretending the human review time is free.
Not after ignoring integration, quality checks, security, and rework.
After the full cost is included.
That is the point where token spend stops being a scary abstraction and becomes part of workflow economics.
A one-dollar AI call can be expensive if it produces a polished answer that takes a senior person thirty minutes to unwind. A fifty-dollar AI run can be cheap if it helps avoid a bad contract clause, compresses a high-value analysis, or lets a team review more material without lowering quality.
The number only means something when it is attached to the work.
The accounting mistake
The common mistake is treating all AI spend as if it should behave the same way.
That creates two bad decisions.
First, leaders over-police exploration. They ask early experiments to produce production-grade ROI before the organization knows which workflows are worth standardizing.
Second, they under-police production. They let recurring AI usage grow because adoption looks good, even when nobody can say whether the workflow improved.
Both are expensive.
Learning spend should be generous enough to discover value and strict enough to prevent wandering.
Production spend should be boringly accountable.
That distinction matters more as AI becomes cheaper to use and easier to hide. When the cost of generating work falls, the cost of generating unnecessary work falls too. More drafts, more summaries, more prototypes, more reports, more internal artifacts that look useful because they exist.
The firm does not need infinite output.
It needs better work.
The portfolio is the point
Executives should think about AI spend as a portfolio.
Most experiments will not become serious enterprise value. Some will produce local productivity. A small number may change the economics of a workflow enough to pay for the broader portfolio.
The exact percentages will vary. The pattern will not.
There will be noise.
There will be learning.
There will be small wins.
There will be a few workflows that matter.
The job is not to make every token profitable. The job is to find the few workflows where AI changes cost, speed, quality, risk, capacity, or customer outcome, then standardize those before the organization pays to rediscover them everywhere.
This is where the executive posture has to be more precise than optimism or skepticism.
Optimism says: AI will transform everything, so spend.
Skepticism says: prove the ROI before we move.
Neither is good enough.
The better posture is:
Spend to learn. Measure to promote. Cut what does not graduate.
Graduation rules
A useful AI experiment should eventually graduate or die.
It should graduate when there is evidence that the workflow improved and the organization knows the conditions under which the pattern works.
It should die when it produces novelty without changed work, savings without a plan for the saved time, speed with too much rework, or risk that cannot be governed.
The graduation criteria do not need to be complicated.
For a workflow to move from learning spend to production spend, the team should be able to answer:
- What workflow is this changing?
- What is the baseline?
- What improved?
- What got worse?
- What human review is required?
- What data or context is allowed?
- What failure modes did we observe?
- What work stops, shrinks, or changes if this succeeds?
- What is the unit cost after AI, review, and rework are included?
If those questions cannot be answered, the work may still be worth exploring. It is not ready to be counted as ROI.
This is the difference between experimentation and drift.
Experimentation has a learning question.
Drift has a tool bill.
The comfort comes from the system
Executives do not need to feel comfortable because someone promised AI will pay off.
They should feel comfortable only if the organization has a system for separating learning from production, promoting what works, and cutting what does not.
That system does not have to be heavy. In fact, if it is too heavy, people will route around it.
But it has to be real.
There should be budgets for exploration. There should be thresholds for production. There should be workflow-level measurement. There should be review standards. There should be a way to stop paying for work that only looks productive because AI made it cheaper to produce.
Token spend is not the investment.
Changed work is the investment.
Tokens are the meter running while the organization finds out whether the work can change.