AI Doesn’t Re-Define What True Analysis Requires Part I

Real analysis requires more than just classifying data into a simple hierarchy, rolling up some numbers, and generating a report or two. It requires more than traditional BI tools which, when all is said and done, don’t do much more than that.

Going back to the beginning when we first defined the New Horizons (Part 1 and Part 2), we find that modern AI still doesn’t always support even the foundations that the Spend Master Eric Strovink defined two decades ago.

Two decades ago, in the original series, we defined:

Meta Aggregation

You don’t just need roll-ups and derived dimensions, you also need “cluster dimensions”. The ability to consolidate spend / vendors in groupings in pre-set ranges (such as the number of consulting vendors under 100K, the number between 100K and 500K, and the number getting the big bucks at more than 500K a year). In other words, classic f(x) doesn’t cut it. You need g(f(x)). You need recursive functions. Which can dynamically update in real time not just when the ranges are changed, but when the hierarchy of the source dimensions is altered (or the source data altered).

Visual Crosstabs

The utility of Shneiderman diagrams (or “treemaps”) to display hierarchical information is well known; the treemap is useful because it is visually intuitive. The relative sizes of the rectangles represent the relative magnitude of spending; the colors indicate relative change in spending; and inner rectangles show the breakdown at the next level of the hierarchy.

Now, suppose that rather than the inner rectangles showing a lower level of the same hierarchy, instead they showed the breakdown of spending within another dimension entirely — i.e., a “visual crosstab.” The visual crosstab would not only show magnitudes, but trends as well.

Federation

 

One of the most serious limitations of OLAP analysis is the schema structure itself — typically a “star” schema, where a voluminous “fact” or “transaction” file is surrounded by supporting files, or “dimensions.” In the case of spend analysis, dimensions are Supplier, Cost Center, Commodity, and so on; transactions are typically AP records.

There are only certain ways that a dimension file can be linked to transaction files, and it isn’t always clear which file ought to be the transaction file and which files ought to be dimensions. For example, suppose that the transaction file consists of AP transactions, and a dimension file consists of invoice line items. The problem is that the invoice line item file is “moving faster” than the AP file; i.e. for every invoice number that appears in AP, there are multiple invoice lines that match. Which invoice line item should we link to?

There are numerous other examples. But the essential problem is that we have two separate datasets, and we’re trying to join them at the hip. There is an AP dataset, and there is an invoice line item dataset, and never the twain shall meet, except artificially. Even when there is no granularity issue at all, and when one dataset can be normalized or snowflaked such that every matching line item can be joined through from the other, the amount of effort required to set up the index->index->index relationships can be daunting.

Instead, why not create two separate datasets, efficiently and quickly; and then as a final step, federate them together on a common dimension? … Federation represents a key productivity enhancer for dataset creation, as well as a simplification to the dataset building process in general.

But that was just the beginning.

AI Has Not Changed the Meaning of Analysis — It’s Reinforced It

Twenty years ago, we took a stab at defining analysis with the help of the Grand Master himself in Defining Analysis, the spend slayer Eric Strovink who put pen to paper on what analysis in practice really is.

1. Agility

To this day, a lot of analytics applications are still based on ROLAP (real-time OLAP [online analytical processing]). ROLAP is great for queries that fall within the rigid framework it was defined for, but to ask a question outside of that rigid framework, and to get an answer to that question in a reasonable amount of time, the underlying dataset structure must be changed. That takes time, if it can be accomplished at all as it typically requires the agreement of multiple stakeholders.

But data analysis is an ad-hoc process. To paraphrase Sun Tzu, “no canned report survives
first contact with the analyst.”
To perform the queries that support the ad-hoc analysis, the analyst needs to:

  • generate new dimensions
  • change existing dimensional hierarchies
  • map and family new and existing dimensions

and that requires an application that supports real agility.

Especially since the ad-hoc analysis not only requires new dimensions and hierarchies, but new datasets, with new cubes and federations on the fly.

2. Speed

Now, with traditional BI tools, you can theoretically do most of the things a modern spend analysis system can do. Load the data sets, cleanse and categorize with queries, write programs to build reports, and dump data to pivot tables in every Purchaser’s favourite tool (Excel) to get different views But that can take days (or weeks) for even a simple problem.

However, you don’t have days or weeks, and sometimes you don’t even have hours to spare when you need an answer quickly. You need an answer right away, or you need to do the same old, same old or, even worse, trust the AI. So you need a tool that supports speed. Whether that be loading, cleansing and categorization, cube building, derived dimension creation, federation, views, filters, and pivots.

3. Power

Every spend analysis system comes with some built in reporting and claims ad-hoc reporting, which at least allows for built-in report customization, but that is not power. Power is the power you wield as a business user, independent of canned reports supplied by a vendor. The creation of customized complex multi-page analysis, from scratch, should be within your reach with minutes of efforts, without any programming skills or super-human training programs.

In the world of AI, these things still don’t exist.

AIs follow scripts, which can be rather rigid. Even worse, Gen-AI creates a script that may or may not be appropriate before executing it. That’s not agility.

AI is blindingly fast as it’s only limited by the computational speed of the processor(s) you give it, but that’s not speed. To achieve real speed, as per the formal definition, distance must be covered. AI can execute hundreds or thousands of canned analyses and rank potential opportunities, but until they are verified, and the factors AI doesn’t know taken into account, it doesn’t generate real sourcing / procurement opportunities and doesn’t cover the distance. That’s not speed.

AI generates reports as it sees them, not as you want them. Gen-AI interprets your requests based on probabilities, not on realities, and that does not give you the power you need to get the insights you need the way you need them to get the support you need to make real progress. It’s not power, just the illusion thereof.

But even worse, it misses the key point that was implied and now needs to be made explicit.

4. Intelligence

It doesn’t have any. The whole point of agility, speed, and power was to support the intuition and insight of an intelligent human who can figure out new ways of looking at data to find new opportunities. Without intelligence, systems are useless. And in the age of AI, the lack of intelligence reinforces the need, and meaning of, analysis more than ever.

AI Has Not Changed the Psychology of Analysis — It’s Aggravated It

It used to be that data analysis that should be performed was avoided because it carried too much risk for the stakeholder due to the time and effort required.

Using original BI tools, when an analysis to determine a question, even a valid one, was likely going to take longer than X hours, and tie up limited resources on a quest that may or may not yield a return (even a decent five figure one), it had to be abandoned.

In addition, when an analysis to answer a question that could save six (or more) figures a year required a change to the organizational classification hierarchy to accomplish, which could require weeks to months to obtain (since classic BI systems could support one, and only one, schema that supported one, and only one, cube definition, and such a change would require the approval of all stakeholders), the analysis was abandoned because the effort to get the agreement could cost more than the analysis could return.

When analysis was not quick, easy, and always available, it didn’t get done when it should — and that’s any time someone had an idea that might save time, money, or both.

Today we have the same problem, but instead of a lack of analysis being performed, AI is performing too many and overloading the organization with “opportunities” of all shapes and sizes being pushed to the sourcing and procurement teams to pursue. And they are overloaded, with instructions to pursue and capture as many opportunities as possible.

Since most providers who provide AI Analytics provide hybrid solutions (where the Gen-AI is given access to traditional, deterministic, solutions that don’t make simple math mistakes), except when the LLM retrieves (or hallucinates) bad data or decides to skip using the deterministic solution for the calculation, most of the opportunities are real to some extent. And since most organizations don’t have good manual software, they’ll only have time to check a few, and when those come back good enough, the organization, out of time, will assume the rest are good (enough) as well and send them off to the sourcing and procurement team.

Some will pan out less than expected, some won’t, but since they generally won’t lose money (even though they won’t capture the savings they should), the organizational leaders will assume they’re doing well, that a manual analysis wouldn’t do much better, and that the volume of opportunities will make up for the lack of significant returns on the pursuit of any individual opportunity.

This is a problem for two reasons. First, the biggest opportunities will go undetected because the AI is running off of scripts, doing the same analysis over and over again, and without knowing all the variables, or having human intuition, won’t investigate where the real opportunities lie. It won’t see the downstream effects of a political tension that will result in a border closing, war, or strait closing that will jack some prices sky high unless demand is locked in now. Nor will it see the the impacts of too many data centres coming online at the same time as AI crash hits and not advise you to delay locking in long term data centre contracts until that happens.

Secondly, it will make mistakes, and sometimes the sourcing and procurement teams will spend a lot of time, possibly weeks or months, on exercises that result in new agreements that actually cost more money than the organization is paying (on average now) because they were undertaken at the wrong time, for the wrong demand, with the wrong suppliers … when existing contracts were still in place for partial demand that couldn’t be broken (without penalty).

These false opportunities won’t be caught because the manpower isn’t there to verify every opportunity (without the right software designed for rapid manual spend analysis), and, as a result, more manpower will be wasted verifying opportunities than just pursing what the AI spits out. However, the results won’t match the savings that would be achieved if the team only pursued verified opportunities (where you verified a real, significant, opportunity). But the lack of time and resources to verify (since management froze hiring to pay for the worthless AI) means the opportunities chased don’t get verified. When it takes more time to second guess the AI than to follow it blindly, the psychology of analysis is to not do it, just like 20 years ago the psychology of analysis was not to do it.

AI Hasn’t Changed the Value Curve of Spend Analysis — It’s Accelerated the Need for Tools to Address It

In our last post on how AI increases the need for manual spend analysis, we noted that the claims that AI negates the need for manual spend analysis is all lies, damn lies, and statistics (of the worst kind).

AI throws more “opportunities” at you than you can process with spreadsheets and traditional AI tools, and it does so faster than an entire team of analysts working in the basement, and, unlike the opportunities coming from the analysts which are at least plausible, the “opportunities” coming from the AI may be completely hallucinatory and a complete waste of time (compared to the opportunities from the junior analysts that may not be worth the effort but are actually real).

So you need to be able to investigate and verify them very, very quickly — which means you need a tool that can:

  • rapidly suck in a dataset, or multiple
  • allow for rapid (re)-classification and (re)-categorization on the fly, using both pre-defined and on-the-fly rules, as well as Gen-AI for suggestions (with quick acceptance or rejection)
  • allow for the quick construction of cubes (and sub-cubes) and views (and sub-views) to verify spend, variances, and potential opportunities against market prices
  • allow for easy generation of derived dimensions and derived cubes with what-if scenarios
  • allow for easy visualization and cross-comparison of scenarios to identify and confirm true opportunities

And then when the opportunities don’t pan out, you need to be able to quickly

  • replicate the data set and rapidly (re)-classify and (re)-categorize on the fly, using both pre-defined and new on-the-fly rules
  • create new cubes (and sub-cubes), views and sub-views on the recategorized data
  • create new derived dimensions (and derived sub-cubes) based on what-if scenarios
  • cross-compare and visualize the new scenarios to identify real opportunities

Or you will never realize any value from spend analysis, especially if you are using AI. The only difference now vs. 20 years ago when you desperately needed a good spend analysis tool was that 20 years ago you didn’t have any opportunities unless you had the time to look for them. Now, you don’t have any opportunities unless you have the time to dig through all of the fake opportunities generated by AI to weed out the real ones from the hallucinations. The value curve is the same, only the ends have switched positions.

AI Increases the Need for Manual Spend Analysis

Everyone with an AI tool is proclaiming how their tool negates the need for manual spend analysis and, beyond that, how it negates the need for sourcing (because their Agentic solution can create the event and execute it automatically), and how it can even handle the contract generation, negotiations, and signing. How it will save you so much time and money that you will never need to do manual spend analysis again.

And it’s all lies, damn lies, and statistics — of the worst kind — since it usually involves probabilistic Gen-AI that often hallucinates market conditions, best practice, and true innovation.

It’s true that an agentic solution can run pre-packaged spend analysis against all of your spend once it has been classified, identify the products and categories with the greatest variance, compare the average price paid to current market prices, suggest opportunities, and push the products into the sourcing tool (and invoke the sourcing agent).

But here are the problems:

  • even assuming it’s ranking opportunities by the savings today, that doesn’t mean it’s the greatest savings tomorrow; there could be contracts in place, declining demand, or rising inflation in the category
  • similarly, products or categories passed over could be the greatest opportunities due to increasing demand, lack of long term contracts, or stagnation in pricing
  • without verification, the “market prices” could be quite off actuals as it could be “public” pricing only and private is quite cheaper, especially with volume discounts
  • the opportunities are not real unless they are with suppliers you can actually source from (and auto-identification is not actual product verification)
  • and just getting auto-responses to auto-requests doesn’t guarantee you get the responses from the suppliers with the best products and prices

And that’s just the beginning.

Here are the missing pieces.

  • it’s working on a predefined categorization (yours, the vendor’s, or, even worse, a random AI one); sometimes the best opportunities come from a categorical reorganization that allows you to bundle the right mix of products for the suppliers you’re inviting
  • suppliers don’t auto-bid in response to auto-RFI requests, vendors do — and for strategic categories, you need strategic suppliers
  • you can use agentic CLM to assemble clauses or send and retrieve documents and even do an initial analysis, but since the review is based on Gen-AI LLMs that can hallucinate as often as not, a review can be helpful, but it’s not a real review
  • the data ingested for product matchings may or may not be accurate, especially since suppliers or distributors can put want they want on their site, or, more importantly, leave out key details

It’s great to use AI to find potential opportunities, but you can’t use AI to verify, and you definitely can’t use AI to capture them. Manual spend analysis has just become more important than ever, as has a tool that allows you to do it as efficiently as AI appears to. Fortunately, a such tools do exist — although you may have to work hard to find one.