AI Doesn’t Change the Fundamentals of Spend Analysis

Twenty years ago, we wrote a post that said, so you want to do spend analysis?

And in that post, we noted that you want to get on with it, but you’re not sure where, or how, to start. So we provided you with a step-by-step process you can use to get your spend analysis effort under way and keep more of those corporate dollars in the corporate coffers, where they belong. And, after 20 years, the steps remain, more or less, the same. Because the song hasn’t changed. Only the noise by those trying to sell you fancy tools (that don’t work) has gotten louder in their attempts to drown the song out.

1. Locate Your Data

You need to identify where all of the source data is. Not the data warehouse or data lake you’re pointed at or someone’s random, monthly data dump, but the source data. That’s because it’s not likely that all of the data you need will be in the warehouse, lake, or dump because that likely consists just of the ERP or AP system, and we know that no single system contains all the data you need. Spend will be split between the ERP, AP, and T&E. Supplier information in the ERP, SXM, and Risk Management Systems. Contracts, and key data, in the CLM and Legal Systems. And so on.

Make sure you know where all the source data you need is, what warehouses and lakes it is pushed to, on what schedule, and which of these you can use. If you can work with system owners to get the missing data pushed to central locations, or if you will need to retrieve it directly.

2. Create a starting taxonomy

You can’t integrate the data otherwise. As long as the taxonomy integrates and federates all of the data, it will be sufficient. Remember that a modern spend analysis system supports multiple cubes, with inheritance, that each cube can have its own hierarchy, and that cubes and sub-cubes can be created on the fly.

3. Centralized the Data

Centralize the data into a single (virtual) starting cube that everything can work off of. Do this by defining the master transaction records, absorb all of the related information, and build the initial (virtual) cube(s) that every (derived) cube will work off of. (By virtual, we mean that the system can pull data on demand as needed with predefined definitions, it doesn’t have to be stored in a single cube in a single warehouse or lake.)

4. Family the Data

In just about any ERP or enterprise system, there are multiple entries for every supplier, product, or other entity. Make sure these are familied into one instance in the master data set.

5. Map the Data

As the data gets pulled in, map to the starting taxonomy so that the data is ready for use.

6. Pick the Low Hanging Fruit

Always start with the easy opportunities — today those might be the highest ranked opportunities spit out by the AI, the ones from the “start here” guide (supplier rationalization, bypass spend identification, growing, off-contract, categories, etc.). A few quick wins will build up support for a real, manual, spend analysis effort — which is where the real results will come from.

7. Make sure you get the right tool!

Remember, the tool, which should have all of the features defined in our previous entry in this series, must be one that gives you the agility, speed, and power you need to apply your intelligence to the problem at hand.

8. Get a good consulting mentor

You want someone to teach you not only how to do spend analysis, but how to approach and even think about spend analysis so that you can translate your instincts into analysis with quick, pinpoint, accuracy and answer your queries faster and better than any AI ever could. That’s not your run of the mill consultant who just executes the standard slate of analysis and applies the AI tools in their stables, that’s a real analyst who understands what analysis truly is and is willing to teach you. They are few and far between, and may not come cheap due to high demand, but once you learn, you won’t need them anymore. And what you spend up front will be insignificant with respect to the opportunities real analysis can uncover.

AI Doesn’t Re-Define What True Analysis Requires Part II

Real analysis requires more than just classifying data into a simple hierarchy, rolling up some numbers, and generating a report or two. It requires more than traditional BI tools which, when all is said and done, don’t do much more than that.

Going back to the beginning when we first defined the New Horizons (Part 1 and Part 2), we find that modern AI still doesn’t always support even the foundations that the Spend Master Eric Strovink defined two decades ago.

But that was just the beginning. Today, we also need:

Cube Reverse Engineering and Easy Rule Construction

If the analyst has an existing cube, then the spend analysis system should be capable of extracting all of the rules flawlessly (as well as detecting the classification errors) and presenting them to the user for easy review. If not, then the system should be able to extract all of the rules used for data organization in the underlying systems.

Moreover, for data that still remains unclassified, it should be easy to use multiple AI techniques to generate suggested rules that can be easily accepted, modified, or rejected.

Reusable Cube Template Definitions

That includes derived, aggregated, and federated dimensions as well as pre-defined views that can be easily mapped to existing data sets, edited, and modified to an analyst’s liking. A user should never have to start from scratch when they did a similar analysis before. And tweaking should be as few clicks as possible.

It should also allow for pre-defined cleansing and categorization rules, easy tweaking thereof, and easy integration of Gen-AI for mapping rule suggestions when the built in capabilities are insufficient.

Reusable Filters

It’s not just cubes that need to be reusable, but filters that allow for customized drill-downs into federated hierarchies — it should be super simple to extract, replicated, modify, and save filters that customized cubes and views to exactly what the user needs to see when they need to see it.

Inheritance

By now, you know the big problem with spend analysis is one-cube with one hierarchy based on one schema. A hierarchy defined (on a schema defined) by consensus satisfies no one. A hierarchy defined by an analyst satisfies only the analyst, and only for the problems she is working on at the time.

As a result, in an average organization, whenever analysis needs to be done, a copy of the data is made, another cube is built, another customization performed, another report printed off, another decision made, and when it comes time to revisit the analysis, or rerun it with updated data, no one knows which of the many data sets is the right one to start with, or even where the most recent data from the source apps got pushed to. Because everything is disconnected, analysis makes a mess.

But if the spend analysis application supports inheritance, where cubes can be derived from existing source cubes, and where data updates are automatically propagated down, but not up, then there only needs to be one copy of the source data, and as long as that source is kept up to date, all of the sub-cubes can be kept up to date as well. They don’t need to be tracked, or even maintained when their usefulness has come to an end, as they are just temporary instances a user can create to answer specific questions, where they can manipulate the data, and hierarchy, to their hearts content without impacting any other user in the organization or corrupting the data.

Because they purport to eliminate the need for these capabilities with their advanced capabilities, most AI systems just don’t provide most of these capabilities. But these are the capabilities that true analysis requires.

AI Doesn’t Re-Define What True Analysis Requires Part I

Real analysis requires more than just classifying data into a simple hierarchy, rolling up some numbers, and generating a report or two. It requires more than traditional BI tools which, when all is said and done, don’t do much more than that.

Going back to the beginning when we first defined the New Horizons (Part 1 and Part 2), we find that modern AI still doesn’t always support even the foundations that the Spend Master Eric Strovink defined two decades ago.

Two decades ago, in the original series, we defined:

Meta Aggregation

You don’t just need roll-ups and derived dimensions, you also need “cluster dimensions”. The ability to consolidate spend / vendors in groupings in pre-set ranges (such as the number of consulting vendors under 100K, the number between 100K and 500K, and the number getting the big bucks at more than 500K a year). In other words, classic f(x) doesn’t cut it. You need g(f(x)). You need recursive functions. Which can dynamically update in real time not just when the ranges are changed, but when the hierarchy of the source dimensions is altered (or the source data altered).

Visual Crosstabs

The utility of Shneiderman diagrams (or “treemaps”) to display hierarchical information is well known; the treemap is useful because it is visually intuitive. The relative sizes of the rectangles represent the relative magnitude of spending; the colors indicate relative change in spending; and inner rectangles show the breakdown at the next level of the hierarchy.

Now, suppose that rather than the inner rectangles showing a lower level of the same hierarchy, instead they showed the breakdown of spending within another dimension entirely — i.e., a “visual crosstab.” The visual crosstab would not only show magnitudes, but trends as well.

Federation

 

One of the most serious limitations of OLAP analysis is the schema structure itself — typically a “star” schema, where a voluminous “fact” or “transaction” file is surrounded by supporting files, or “dimensions.” In the case of spend analysis, dimensions are Supplier, Cost Center, Commodity, and so on; transactions are typically AP records.

There are only certain ways that a dimension file can be linked to transaction files, and it isn’t always clear which file ought to be the transaction file and which files ought to be dimensions. For example, suppose that the transaction file consists of AP transactions, and a dimension file consists of invoice line items. The problem is that the invoice line item file is “moving faster” than the AP file; i.e. for every invoice number that appears in AP, there are multiple invoice lines that match. Which invoice line item should we link to?

There are numerous other examples. But the essential problem is that we have two separate datasets, and we’re trying to join them at the hip. There is an AP dataset, and there is an invoice line item dataset, and never the twain shall meet, except artificially. Even when there is no granularity issue at all, and when one dataset can be normalized or snowflaked such that every matching line item can be joined through from the other, the amount of effort required to set up the index->index->index relationships can be daunting.

Instead, why not create two separate datasets, efficiently and quickly; and then as a final step, federate them together on a common dimension? … Federation represents a key productivity enhancer for dataset creation, as well as a simplification to the dataset building process in general.

But that was just the beginning.

AI Has Not Changed the Meaning of Analysis — It’s Reinforced It

Twenty years ago, we took a stab at defining analysis with the help of the Grand Master himself in Defining Analysis, the spend slayer Eric Strovink who put pen to paper on what analysis in practice really is.

1. Agility

To this day, a lot of analytics applications are still based on ROLAP (real-time OLAP [online analytical processing]). ROLAP is great for queries that fall within the rigid framework it was defined for, but to ask a question outside of that rigid framework, and to get an answer to that question in a reasonable amount of time, the underlying dataset structure must be changed. That takes time, if it can be accomplished at all as it typically requires the agreement of multiple stakeholders.

But data analysis is an ad-hoc process. To paraphrase Sun Tzu, “no canned report survives
first contact with the analyst.”
To perform the queries that support the ad-hoc analysis, the analyst needs to:

  • generate new dimensions
  • change existing dimensional hierarchies
  • map and family new and existing dimensions

and that requires an application that supports real agility.

Especially since the ad-hoc analysis not only requires new dimensions and hierarchies, but new datasets, with new cubes and federations on the fly.

2. Speed

Now, with traditional BI tools, you can theoretically do most of the things a modern spend analysis system can do. Load the data sets, cleanse and categorize with queries, write programs to build reports, and dump data to pivot tables in every Purchaser’s favourite tool (Excel) to get different views But that can take days (or weeks) for even a simple problem.

However, you don’t have days or weeks, and sometimes you don’t even have hours to spare when you need an answer quickly. You need an answer right away, or you need to do the same old, same old or, even worse, trust the AI. So you need a tool that supports speed. Whether that be loading, cleansing and categorization, cube building, derived dimension creation, federation, views, filters, and pivots.

3. Power

Every spend analysis system comes with some built in reporting and claims ad-hoc reporting, which at least allows for built-in report customization, but that is not power. Power is the power you wield as a business user, independent of canned reports supplied by a vendor. The creation of customized complex multi-page analysis, from scratch, should be within your reach with minutes of efforts, without any programming skills or super-human training programs.

In the world of AI, these things still don’t exist.

AIs follow scripts, which can be rather rigid. Even worse, Gen-AI creates a script that may or may not be appropriate before executing it. That’s not agility.

AI is blindingly fast as it’s only limited by the computational speed of the processor(s) you give it, but that’s not speed. To achieve real speed, as per the formal definition, distance must be covered. AI can execute hundreds or thousands of canned analyses and rank potential opportunities, but until they are verified, and the factors AI doesn’t know taken into account, it doesn’t generate real sourcing / procurement opportunities and doesn’t cover the distance. That’s not speed.

AI generates reports as it sees them, not as you want them. Gen-AI interprets your requests based on probabilities, not on realities, and that does not give you the power you need to get the insights you need the way you need them to get the support you need to make real progress. It’s not power, just the illusion thereof.

But even worse, it misses the key point that was implied and now needs to be made explicit.

4. Intelligence

It doesn’t have any. The whole point of agility, speed, and power was to support the intuition and insight of an intelligent human who can figure out new ways of looking at data to find new opportunities. Without intelligence, systems are useless. And in the age of AI, the lack of intelligence reinforces the need, and meaning of, analysis more than ever.

AI Has Not Changed the Psychology of Analysis — It’s Aggravated It

It used to be that data analysis that should be performed was avoided because it carried too much risk for the stakeholder due to the time and effort required.

Using original BI tools, when an analysis to determine a question, even a valid one, was likely going to take longer than X hours, and tie up limited resources on a quest that may or may not yield a return (even a decent five figure one), it had to be abandoned.

In addition, when an analysis to answer a question that could save six (or more) figures a year required a change to the organizational classification hierarchy to accomplish, which could require weeks to months to obtain (since classic BI systems could support one, and only one, schema that supported one, and only one, cube definition, and such a change would require the approval of all stakeholders), the analysis was abandoned because the effort to get the agreement could cost more than the analysis could return.

When analysis was not quick, easy, and always available, it didn’t get done when it should — and that’s any time someone had an idea that might save time, money, or both.

Today we have the same problem, but instead of a lack of analysis being performed, AI is performing too many and overloading the organization with “opportunities” of all shapes and sizes being pushed to the sourcing and procurement teams to pursue. And they are overloaded, with instructions to pursue and capture as many opportunities as possible.

Since most providers who provide AI Analytics provide hybrid solutions (where the Gen-AI is given access to traditional, deterministic, solutions that don’t make simple math mistakes), except when the LLM retrieves (or hallucinates) bad data or decides to skip using the deterministic solution for the calculation, most of the opportunities are real to some extent. And since most organizations don’t have good manual software, they’ll only have time to check a few, and when those come back good enough, the organization, out of time, will assume the rest are good (enough) as well and send them off to the sourcing and procurement team.

Some will pan out less than expected, some won’t, but since they generally won’t lose money (even though they won’t capture the savings they should), the organizational leaders will assume they’re doing well, that a manual analysis wouldn’t do much better, and that the volume of opportunities will make up for the lack of significant returns on the pursuit of any individual opportunity.

This is a problem for two reasons. First, the biggest opportunities will go undetected because the AI is running off of scripts, doing the same analysis over and over again, and without knowing all the variables, or having human intuition, won’t investigate where the real opportunities lie. It won’t see the downstream effects of a political tension that will result in a border closing, war, or strait closing that will jack some prices sky high unless demand is locked in now. Nor will it see the the impacts of too many data centres coming online at the same time as AI crash hits and not advise you to delay locking in long term data centre contracts until that happens.

Secondly, it will make mistakes, and sometimes the sourcing and procurement teams will spend a lot of time, possibly weeks or months, on exercises that result in new agreements that actually cost more money than the organization is paying (on average now) because they were undertaken at the wrong time, for the wrong demand, with the wrong suppliers … when existing contracts were still in place for partial demand that couldn’t be broken (without penalty).

These false opportunities won’t be caught because the manpower isn’t there to verify every opportunity (without the right software designed for rapid manual spend analysis), and, as a result, more manpower will be wasted verifying opportunities than just pursing what the AI spits out. However, the results won’t match the savings that would be achieved if the team only pursued verified opportunities (where you verified a real, significant, opportunity). But the lack of time and resources to verify (since management froze hiring to pay for the worthless AI) means the opportunities chased don’t get verified. When it takes more time to second guess the AI than to follow it blindly, the psychology of analysis is to not do it, just like 20 years ago the psychology of analysis was not to do it.