Skip to content
       

Blog

How to Tell If Your AI Is Actually Working, or Just Busy

How to Tell If Your AI Is Actually Working, or Just Busy

Here is one of the more uncomfortable findings in enterprise technology right now, and it is worth sitting with before a property company approves another dollar of AI spend. A large share of organizations believe they are getting value from AI, and only a tiny fraction can actually prove it. An MIT study in 2025 found that roughly 95% of enterprises saw zero measurable return on their AI investments, and Deloitte's research found that fewer than one in three organizations can measure their AI ROI with confidence despite widespread deployment. That is not a small measurement gap. That is most companies operating on a feeling.

The feeling is not baseless, exactly. Things are visibly happening. Owner reports get written faster, tenant emails get drafted more quickly, maintenance summaries get produced in higher volume, dashboards fill with green. The operation is unmistakably busier with AI. The problem is that busy and working are not the same thing, and the entire discipline of knowing which one you have is the difference between an AI investment that pays off and one that merely feels like it does. Learning to tell them apart is arguably the most important AI skill a leadership team can develop, and almost no one is teaching it, because most of the people talking about AI have an interest in you not looking too closely.

The default state is theater

The scale of the problem is documented, not anecdotal. Set the near-total absence of measurable return next to the enthusiasm and the spending, which runs into the hundreds of billions of dollars globally, and you have the defining tension of this moment: enormous investment, near-universal belief in its value, and almost no ability to prove that value exists. The pattern holds even among the biggest adopters. McKinsey's research found that while 88% of organizations use AI in at least one business function, only about 6% qualify as high performers, meaning they can attribute a meaningful share of company earnings to AI, and the telling detail is that those high performers are not using more sophisticated AI than everyone else. They are measuring it differently.

The industry has a name for what fills the gap between spending and proof, and it is a good one: productivity theater. It is the appearance of value, activity that looks like progress, tracked with metrics that are always large and always impressive, standing in for evidence that the business actually improved. The tell is that when the CFO finally asks for the return, the team produces activity metrics, emails drafted, documents generated, hours ostensibly saved, that do not map to any financial outcome. Everyone feels more productive. No one can point to a number on the income statement that moved. That is not a fraud anyone committed on purpose. It is the natural default when no one built the measurement in first, and it is where most organizations sit right now whether they know it or not.

Why the illusion is so durable

Productivity theater persists because nearly every incentive in the system props it up, which is why seeing through it requires deliberate effort rather than good intentions.

Vendors encourage activity metrics because they flatter the product. "Time saved" and "tasks automated" are always big numbers, and the vendor has no reason to help you check whether that saved time turned into anything the business can bank. Managers who championed an AI rollout need it to have worked, and "our team saves ten hours a week" is a compelling line in a performance review even if those ten hours were never redirected to anything of value. And there is a genuine psychological effect underneath both: people using a new tool feel more productive almost regardless of whether their actual output changed, an echo of the old Hawthorne effect where being observed and doing something new produces a temporary lift in perceived performance that early surveys faithfully capture and that fades once the novelty does.

So the enthusiasm is real, the activity is real, and the metrics are real. What is missing is the one link that matters: proof that any of it changed a business outcome. And because the incentives all point away from checking that link, the illusion is stable. It does not collapse on its own. Someone has to decide to look.

The honest part: not everything shows up in a quarter

It would be a mistake to swing to the opposite error, which is a kind of ROI fundamentalism that dismisses anything not immediately visible on this quarter's P&L. Some genuine AI value is real but slow, or real but indirect, and demanding instant financial proof can kill things that would have paid off with time. The internet looked like a poor investment in 1995 by the standard of immediate corporate profit. Some of AI's most real value shows up as cost avoidance rather than revenue, which is quieter and easy to miss in a top-line-focused review. Deloitte's own research notes that AI frequently delivers outcomes that matter but are genuinely hard to monetize, and that many organizations reach satisfactory returns over a two-to-four-year horizon rather than in a quarter. Insisting that everything prove itself in ninety days would be its own kind of blindness.

But there is a sharp difference between "this value is real but will take time to show up in a measurable way" and "we have no way of knowing whether this is working and we are not trying to find out." The first is a legitimate, honest position, provided you name the outcome you eventually expect and the horizon on which to check for it. The second is theater.

The ROI test: Can you state, in advance, the specific business outcome that must become true for this to count as working, and the exact date you will measure it? If yes, you're measuring. If no, you're hoping.

The difference between busy and working

The discipline that separates them is not complicated, which makes its rarity all the more telling. It has three parts, and skipping any one of them collapses the whole thing back into theater.

First, baseline before you deploy. You cannot measure improvement against a number you never recorded, and the single most common failure is rolling out AI without capturing what performance looked like beforehand. If you did not measure the close time, the error rate, the cost per unit, or the vacancy-days before the AI arrived, you have permanently lost the ability to prove the AI changed them. The baseline is not paperwork. It is the only thing that makes a later claim of improvement more than an assertion.

Second, measure outcomes, not activity. Activity metrics count inputs, prompts submitted, licenses activated, hours ostensibly saved. Outcome metrics count results, did the error rate fall, did the cycle time shorten, did cost per unit drop, did the thing the business actually cares about move. Only outcome metrics produce a return a CFO can read, and the research bears out that this is not a stylistic preference: Deloitte's AI Institute found that organizations measuring AI with outcome metrics are about 2.4 times more likely to report strong ROI than those tracking only activity. That is almost certainly because the outcome-measurers are the ones actually finding out whether it works, and then fixing or killing what does not.

Third, control for everything else that changed. If your close got faster the same quarter you also hired two people and cleaned up a process, you cannot hand the AI the credit without accounting for the rest. Real measurement isolates the AI's contribution from the other things happening at the same time, which is unglamorous and is exactly the step that separates a defensible ROI claim from a flattering coincidence. "It got better after we adopted AI" is not evidence the AI did it. It is evidence you should go find out.

Put those three together and you have the whole discipline: know where you started, measure whether the business actually changed, and prove the AI is why. An organization doing all three can tell you whether its AI is working. An organization doing none of them can only tell you whether it feels busy, and the research suggests that describes the overwhelming majority. This is also why the underlying data matters so much: an AI running on one clean, connected record is one whose effect on close time or error rate you can actually isolate and measure, while one stitched across disagreeing systems produces results you can neither trust nor attribute. The point is not to be cynical about AI. It is to hold it to the same standard as any other investment, so that the money follows what is genuinely working rather than what is merely, expensively, in motion.

FAQs

Q1. Why can't most companies prove their AI is working?
Because they measure activity instead of outcomes and never established a baseline before deploying. The result is a lot of impressive-looking usage data that does not connect to any business result, so when someone asks for the financial return, there is nothing to point to. Research this year found the large majority of enterprises could show no measurable return, and fewer than one in three can measure their AI ROI with confidence.

Q2. What is "productivity theater"?
It is the appearance of value standing in for evidence of it: activity that looks like progress, tracked with large, flattering metrics, without proof that any business outcome improved. It is the natural default when measurement is not built in from the start, and it persists because vendors, internal champions, and even a psychological novelty effect all push toward counting activity rather than results.

Q3. Isn't some AI value genuinely hard to measure?
Yes, and that is a real and important caveat. Some value is slow, indirect, or shows up as cost avoidance rather than revenue, and demanding instant financial proof can kill things that would eventually pay off, with much AI value arriving over a two-to-four-year horizon. The distinction that matters is whether you can name, in advance, what outcome you expect and when you will check for it. Legitimate patience names its target; theater simply avoids the question.

Q4. What is the difference between activity metrics and outcome metrics?
Activity metrics count what people do with the tool: prompts submitted, licenses activated, hours reportedly saved. Outcome metrics count whether the business changed: error rates, cycle times, cost per unit, vacancy days, conversion. Only outcome metrics produce a return leadership can actually act on, and organizations that track them are considerably more likely to report and prove real ROI.

Q5. Why does a baseline matter so much?
Because improvement is meaningless without a before to compare against. If you did not record the close time, error rate, or cost before the AI arrived, you have lost the ability to prove the AI changed them, and you are left asserting improvement rather than demonstrating it. Capturing the baseline before deployment is the cheapest, highest-value measurement step, and the one most often skipped.

Q6. How do we know the AI caused an improvement and not something else?
By controlling for the other changes happening at the same time. If performance improved in a period when you also added staff or reworked a process, the AI cannot simply be handed the credit. Isolating the AI's specific contribution is what turns a flattering correlation into a defensible ROI claim, and skipping it is how coincidences get reported as wins.

Q7. Does this mean we should be skeptical of all AI ROI claims?
Skeptical in a specific, constructive way: ask what baseline the claim rests on, whether it measures an outcome or just activity, and whether other simultaneous changes were controlled for. A claim that survives those three questions is credible. One that cannot answer them is theater, regardless of how large or confident the number sounds. The goal is not cynicism but rigor.

Q8. What is the first step to measuring our AI honestly?
Before the next deployment, record the current performance of whatever the AI is meant to improve, and define the specific outcome that would count as success and the date you will check it. That single act, baselining and naming the target in advance, converts a future of hopeful assertions into an actual test, and it is the step that most cleanly separates organizations that know whether their AI works from those that only feel that it does.