Skip to content
       

Blog

How to Tell If Your AI Is Actually Working, or Just Busy

How to Tell If Your AI Is Actually Working, or Just Busy

Here is one of the more uncomfortable findings in enterprise technology right now, and it is worth sitting with before you approve another dollar of AI spend. A large share of organizations believe they are getting value from AI, and only a tiny fraction can actually prove it. One Harvard Business Review analysis this year put the gap starkly: roughly 79% of organizations believe they are getting value from AI, while only about 5% can demonstrate it financially. That is not a small measurement gap. That is most companies operating on a feeling.

The feeling is not baseless, exactly. Things are visibly happening. Reports get written faster, emails get drafted more quickly, code gets produced in higher volume, dashboards fill with green. The organization is unmistakably busier with AI. The problem is that busy and working are not the same thing, and the entire discipline of knowing which one you have is the difference between an AI investment that pays off and one that merely feels like it does. Learning to tell them apart is arguably the most important AI skill a leadership team can develop, and almost no one is teaching it, because most of the people talking about AI have an interest in you not looking too closely.

The default state is theater

The scale of the problem is documented, not anecdotal. A widely cited MIT study in 2025 found that roughly 95% of enterprises saw zero measurable return on their AI investments. Zero. Set that next to the enthusiasm and the spending, which Gartner projects will exceed $300 billion globally in 2026, and you have the defining tension of this moment: enormous investment, near-universal belief in its value, and almost no ability to prove that value exists.

The industry has a name for what fills that gap, and it is a good one: productivity theater. It is the appearance of value, activity that looks like progress, tracked with metrics that are always large and always impressive, standing in for evidence that the business actually improved. The tell is that when the CFO finally asks for the return, the team produces activity metrics, emails drafted, documents generated, hours ostensibly saved, that do not map to any financial outcome. Everyone feels more productive. No one can point to a number on the income statement that moved. That is not a fraud anyone committed on purpose. It is the natural default when no one built the measurement in first, and it is where most organizations sit right now whether they know it or not.

Why the illusion is so durable

Productivity theater persists because nearly every incentive in the system props it up, which is why seeing through it requires deliberate effort rather than good intentions.

Vendors encourage activity metrics because they flatter the product. "Time saved" and "tasks automated" are always big numbers, and the vendor has no reason to help you check whether that saved time turned into anything the business can bank. Managers who championed an AI rollout need it to have worked, and "our team saves ten hours a week" is a compelling line in a performance review even if those ten hours were never redirected to anything of value. And there is a genuine psychological effect underneath both: people using a new tool feel more productive almost regardless of whether their actual output changed, an echo of the old Hawthorne effect where being observed and doing something new produces a temporary lift in perceived performance that early surveys faithfully capture and that fades once the novelty does.

So the enthusiasm is real, the activity is real, and the metrics are real. What is missing is the one link that matters: proof that any of it changed a business outcome. And because the incentives all point away from checking that link, the illusion is stable. It does not collapse on its own. Someone has to decide to look.

The honest part: not everything shows up in a quarter

It would be a mistake to swing to the opposite error, which is a kind of ROI fundamentalism that dismisses anything not immediately visible on this quarter's P&L. Some genuine AI value is real but slow, or real but indirect, and demanding instant financial proof can kill things that would have paid off with time. The internet looked like a poor investment in 1995 by the standard of immediate corporate profit. Some of AI's most real value shows up as cost avoidance rather than revenue, which is quieter and easy to miss in a top-line-focused review. And some of it is capability building that pays off later. Insisting that everything prove itself in ninety days would be its own kind of blindness.

But there is a sharp difference between "this value is real but will take time to show up in a measurable way" and "we have no way of knowing whether this is working and we are not trying to find out." The first is a legitimate, honest position, provided you name the outcome you eventually expect and the horizon on which to check for it. The second is theater.

The ROI test: Can you state, in advance, the specific business outcome that must become true for this to count as working, and the exact date you will measure it? If yes, you're measuring. If no, you're hoping.

The difference between busy and working

The discipline that separates them is not complicated, which makes its rarity all the more telling. It has three parts, and skipping any one of them collapses the whole thing back into theater.

First, baseline before you deploy. You cannot measure improvement against a number you never recorded, and the single most common failure is rolling out AI without capturing what performance looked like beforehand. If you did not measure the close time, the error rate, the cost per unit, or the cycle time before the AI arrived, you have permanently lost the ability to prove the AI changed them. The baseline is not paperwork. It is the only thing that makes a later claim of improvement more than an assertion.

Second, measure outcomes, not activity. Activity metrics count inputs, prompts submitted, licenses activated, hours ostensibly saved. Outcome metrics count results, did the error rate fall, did the cycle time shorten, did cost per unit drop, did the thing the business actually cares about move. Only outcome metrics produce a return a CFO can read, and the research bears out that this is not a stylistic preference: Deloitte's AI Institute found that organizations measuring AI with outcome metrics are about 2.4 times more likely to report strong ROI than those tracking only activity. That is almost certainly because the outcome-measurers are the ones actually finding out whether it works, and then fixing or killing what does not.

Third, control for everything else that changed. If your close got faster the same quarter you also hired two people and cleaned up a process, you cannot hand the AI the credit without accounting for the rest. Real measurement isolates the AI's contribution from the other things happening at the same time, which is unglamorous and is exactly the step that separates a defensible ROI claim from a flattering coincidence. "It got better after we adopted AI" is not evidence the AI did it. It is evidence you should go find out.

Put those three together and you have the whole discipline: know where you started, measure whether the business actually changed, and prove the AI is why. An organization doing all three can tell you whether its AI is working. An organization doing none of them can only tell you whether it feels busy, and the research suggests that describes the overwhelming majority. The point is not to be cynical about AI. It is to hold it to the same standard as any other investment, so that the money follows what is genuinely working rather than what is merely, expensively, in motion.

FAQs

Q1. Why can't most companies prove their AI is working?
Because they measure activity instead of outcomes and never established a baseline before deploying. The result is a lot of impressive-looking usage data that does not connect to any business result, so when someone asks for the financial return, there is nothing to point to. Research this year found most organizations believe they are getting value while only a small fraction can actually demonstrate it.

Q2. What is "productivity theater"?
It is the appearance of value standing in for evidence of it: activity that looks like progress, tracked with large, flattering metrics, without proof that any business outcome improved. It is the natural default when measurement is not built in from the start, and it persists because vendors, internal champions, and even a psychological novelty effect all push toward counting activity rather than results.

Q3. Isn't some AI value genuinely hard to measure?
Yes, and that is a real and important caveat. Some value is slow, indirect, or shows up as cost avoidance rather than revenue, and demanding instant financial proof can kill things that would eventually pay off. The distinction that matters is whether you can name, in advance, what outcome you expect and when you will check for it. Legitimate patience names its target; theater simply avoids the question.

Q4. What is the difference between activity metrics and outcome metrics?
Activity metrics count what people do with the tool: prompts submitted, licenses activated, hours reportedly saved. Outcome metrics count whether the business changed: error rates, cycle times, cost per unit, conversion. Only outcome metrics produce a return leadership can actually act on, and organizations that track them are considerably more likely to report and prove real ROI.

Q5. Why does a baseline matter so much?
Because improvement is meaningless without a before to compare against. If you did not record the close time, error rate, or cost before the AI arrived, you have lost the ability to prove the AI changed them, and you are left asserting improvement rather than demonstrating it. Capturing the baseline before deployment is the cheapest, highest-value measurement step, and the one most often skipped.

Q6. How do we know the AI caused an improvement and not something else?
By controlling for the other changes happening at the same time. If performance improved in a period when you also added staff or reworked a process, the AI cannot simply be handed the credit. Isolating the AI's specific contribution is what turns a flattering correlation into a defensible ROI claim, and skipping it is how coincidences get reported as wins.

Q7. Does this mean we should be skeptical of all AI ROI claims?
Skeptical in a specific, constructive way: ask what baseline the claim rests on, whether it measures an outcome or just activity, and whether other simultaneous changes were controlled for. A claim that survives those three questions is credible. One that cannot answer them is theater, regardless of how large or confident the number sounds. The goal is not cynicism but rigor.

Q8. What is the first step to measuring our AI honestly?
Before the next deployment, record the current performance of whatever the AI is meant to improve, and define the specific outcome that would count as success and the date you will check it. That single act, baselining and naming the target in advance, converts a future of hopeful assertions into an actual test, and it is the step that most cleanly separates organizations that know whether their AI works from those that only feel that it does.