Skip to content
       

Blog

The Spreadsheet That Runs Your Company

The Spreadsheet That Runs Your Company

You know the one. It might be the allocation workbook, the distribution waterfall, the owner statement reconciliation that pulls three exports together, or the model that turns your reported numbers into the version the lender sees. It has a single author, a filename with a version suffix that stopped meaning anything years ago, and a set of formulas that only one person fully understands. And a number that appears in your financial statements passes through it every month.

Most organizations regard this as a minor embarrassment, a sign that people are not using the system properly, and periodically resolve to eliminate it. That reaction throws away the most useful information the spreadsheet contains. Its existence is not a failure of discipline. It is precise evidence about your systems, generated by someone who was trying to get their job done, and almost nobody reads it that way.

The risk is measured, not anecdotal

Before the reframe, the risk deserves to be stated properly, because it is more serious than the casual attitude toward spreadsheets suggests.

Research on spreadsheet errors is extensive and consistent. Raymond Panko of the University of Hawaii, whose work is the reference point in this field, synthesized field audits of real operational spreadsheets and found that the more recent audits, using better methodologies, found errors in at least 86% of the spreadsheets they examined. The European Spreadsheet Risks Interest Group, which has coordinated much of this research, puts the figure above 90%, and notes that these errors persist largely because spreadsheets are rarely tested at all.

Panko's conclusions are worth stating in full, because the third one is the dangerous one. Errors are rare on a per-cell basis, but in a large workbook at least one incorrect bottom-line value is very likely. Errors are extremely difficult to detect and correct. And organizations are highly overconfident about the accuracy of their spreadsheets.

That last point undermines the control most companies believe they have. "We check it carefully" is not a meaningful safeguard, because human error detection is substantially worse than human error avoidance. People perform calculations at high accuracy and then find only a fraction of the mistakes that do occur, with detection rates in structured software inspection running well below half. A careful review of a complex workbook is not the control it feels like.

The consequences are not hypothetical. JPMorgan's own internal task force report on the London Whale losses described a value-at-risk model that operated through a series of Excel spreadsheets requiring manual copying between them, containing a formula that divided by a sum where it should have divided by an average, with the effect of understating volatility. A material risk number, produced by a workbook, wrong.

Why the usual response never works

Every few years, an organization declares that critical work will move out of spreadsheets and into the system. It rarely succeeds, and the reason is worth understanding rather than treating as a discipline problem.

A spreadsheet is not the problem. It is somebody's solution to a problem. It exists because the official system could not do something that needed doing, and a competent person solved it anyway, on their own time, using the tool available. Building and maintaining a real workbook is genuine work. Nobody does it recreationally. The existence of a maintained, load-carrying spreadsheet is evidence that someone judged the effort of building and running it to be less painful than the gap in the system.

So banning spreadsheets without addressing the underlying gap does not remove the work. It relocates it, usually somewhere less visible, onto a personal drive or a laptop, where it continues to feed the same numbers with even less oversight. You have not eliminated the risk. You have made it harder to find, which is strictly worse.

Read the inventory as a diagnostic

Here is the reframe worth adopting. Your critical spreadsheets constitute a detailed, accurate, continuously updated map of where your systems fail to meet the business's actual requirements, produced at no cost by the people who know best, and most organizations throw this away as noise.

Each recurring, load-carrying workbook marks a specific location where something was missing. A capability the system does not have. A process that does not match how the work is really done. A question the reporting cannot answer. An exception the configuration does not handle. A calculation that exists in your business but not in your software.

This information is unusually honest, because it reflects revealed preference rather than stated opinion. When you ask people in a requirements workshop what the system needs to do, you get a wish list shaped by what they think is possible and what they remember to mention. When you look at what they actually built a spreadsheet to accomplish, you are looking at a need real enough that someone did unpaid extra work to meet it. That is a considerably stronger signal than anything a survey produces.

Sorting what you find

Not every spreadsheet is telling you the same thing, and the useful discipline is separating three categories rather than treating them as one problem.

Genuine analytical work. Exploratory modeling, one-off analysis, scenario testing, thinking a problem through with numbers. Spreadsheets are excellent at this and always will be. Nothing else lets a competent finance person answer a novel question in an afternoon. Leave these alone entirely.

Gap markers. Recurring workbooks that feed real numbers into real processes. These are the diagnostic material. Each represents a system requirement that nobody ever wrote down, discovered through use rather than through analysis, and it is worth knowing what each one is compensating for.

Shadow systems of record. Workbooks that have quietly become the authoritative version of something, which other people now depend on and other processes now consume. These carry the highest risk, because the organization is treating a file as infrastructure without any of the controls it applies to actual infrastructure, and usually without anyone having decided that this was the arrangement.

The honest part: spreadsheets are good

It would be easy to turn this into an anti-spreadsheet argument, and that position is overstated in ways worth naming.

Spreadsheets remain one of the most useful business tools ever built. They are flexible, immediate, and universally understood, and they let people model something new without a project, a budget, or a developer. An organization that successfully eliminated all spreadsheet use would have made itself considerably slower and less capable of answering questions it had not anticipated.

It is also not true that every identified gap should be closed in the system. Some gaps are genuinely one-off, and building permanent functionality to handle a situation that arises twice a year is over-engineering, which carries its own costs as argued in when good enough is the right call. A spreadsheet is often the correct, proportionate answer to a small, infrequent need, and formalizing it would be waste.

And the error research describes spreadsheets in aggregate. A well-structured, single-purpose, reviewed workbook with a competent owner is a different object from a fifteen-year-old inherited model that four people have modified and nobody has tested. The category is risky. Individual instances vary enormously.

The complication AI adds

There is a timely reason this matters more than it did a few years ago. When an organization deploys AI against its operational data, the AI sees what is in the systems. It does not see the workbook on someone's drive where the actual calculation lives.

So wherever meaningful business logic has migrated into spreadsheets, it is invisible to the tools you are now buying to analyze your business, which means those tools are reasoning about a version of your operation with pieces missing. This compounds the grounding problem covered in why ChatGPT won't fix your property business: it is not only that general models lack access to your data, it is that some of your most important logic is not in the place where any tool would look for it.

AI also makes generating spreadsheets and ad hoc analysis substantially easier, which means the inventory will grow faster than it did when building a model took real effort. The rationing mechanism that limited spreadsheet sprawl was the work involved, and that mechanism is weakening.

Where to start

The practical move is not an inventory of every file in the company, which produces a large document and no action. It is a targeted trace.

Start from your financial statements and work backwards. Which reported numbers pass through a spreadsheet at some point on their way to the statement? That question is answerable in a few conversations, it identifies the workbooks that carry actual consequence, and it usually surprises people, both in how many there are and in which ones turn out to matter.

For each one identified, ask two questions. What system gap is this compensating for, which turns the workbook into information about your requirements rather than a nuisance. And what would happen if it produced a wrong number, which tells you how much control it warrants. Then make a deliberate choice among three options: close the gap in the system, formalize the spreadsheet with version control, review, and a named owner, or consciously accept it as a small risk that is not worth the effort to eliminate. All three are legitimate. What is not legitimate is the current default in most organizations, which is to have never made the choice at all, while a number in the audited financials depends on a file nobody has examined.

FAQs

Q1. How common are errors in operational spreadsheets?
Very. Panko's synthesis of field audits found errors in at least 86% of spreadsheets examined in the more rigorous studies, and EuSpRIG puts the figure above 90%. His most uncomfortable conclusion is that organizations are highly overconfident about the accuracy of their own spreadsheets, meaning the perceived risk is consistently lower than the measured one.

Q2. We review our critical spreadsheets carefully. Isn't that enough?
Careful review is weaker than it feels, because humans are considerably better at avoiding errors than at detecting them. Structured inspection in software development catches well under half of existing defects, and spreadsheet detection rates are comparable. A review provides some assurance but should not be treated as equivalent to a system control with enforced logic and an audit trail.

Q3. Why do spreadsheet elimination programs usually fail?
Because they target the symptom. A spreadsheet exists because the system could not do something and someone solved it anyway, at real personal effort. Removing the tool without addressing the gap relocates the work to a less visible place, often a personal drive, where it continues to produce the same numbers with less oversight. The risk is not reduced, only concealed.

Q4. What does it mean to read spreadsheets as a diagnostic?
It means treating each recurring, load-carrying workbook as evidence of a specific unmet system requirement. Because building and maintaining one is genuine work, its existence reveals a need strong enough to justify unpaid effort, which is a more reliable signal than a requirements workshop produces. The inventory is a map of gaps, generated free by the people closest to the work.

Q5. Should we move all spreadsheet logic into our systems?
No. Some spreadsheet use is genuine analytical work that systems should not replace, and some gaps are infrequent enough that building permanent functionality would be over-engineering. The goal is a deliberate decision for each material workbook, choosing among closing the gap, formalizing the spreadsheet with controls, or knowingly accepting it, rather than having made no decision.

Q6. Which spreadsheets should we worry about most?
Those that have become unofficial systems of record, where a file has quietly become the authoritative version of something and other people and processes now depend on it. These carry infrastructure-level consequence without infrastructure-level controls, and usually nobody ever decided that this should be the arrangement. It happened by accumulation.

Q7. How does this affect our AI plans?
Directly. AI deployed against your operational data sees what is in your systems and cannot see the workbook where a calculation actually lives. Wherever significant logic has migrated into spreadsheets, your AI is reasoning about an incomplete version of your business. AI also makes generating new spreadsheets easier, weakening the effort-based constraint that previously limited how fast the inventory grew.

Q8. What is the most efficient way to find the ones that matter?
Work backwards from your financial statements and identify which reported numbers pass through a spreadsheet on the way. This is answerable in a handful of conversations, concentrates attention on workbooks with real consequence, and avoids the trap of cataloging every file in the organization, which produces a large document and no decisions.