Category Archives: rants

There’s a Reason that AI Talk at (Procure)Tech Events is All Talk and No Substance

A recent rant on LinkedIn noted that, at DPW (and other events), there was “Tons of talk about using AI, not so much about how to procure it, scope it, contract it, measure it, commercial agreements, which suppliers are best for which usage.”

Well, there’s a simple reason for that. And it goes as follows.

It’s the tech-du-jour. More specifically, the hype-du-jour. As a result you can forget any advice on how to:

–> Procure It

No one really knows who has what, or what it actually does, what it’s really worth, how to properly cost it, how to compare offerings and offers. So how can they tell you how to procure it?

–> Scope It

According to the hype it’s your new employee that does everything for every one, despite the fact that it hallucinates more times per second than an LSD junkie does in a lifetime. Without knowing what it does or where it goes, they can’t tell you how to scope it!

–> Contract It / Commercial Agreements

Even the vendors wrapping someone else’s model don’t really know what it does, what they can guarantee, what they can’t, or who is really liable (although courts are starting to make those decisions). So no one knows how to contract it.

–> Measure It

Whereas we have tried-and-true mathematically sound measures for traditional deep neural networks where you can get accuracy ranges with confidence ranges, when it comes to LLMs, no one has a f*ck1ng clue how to measure them. Random tests by random humans judged by random people with random definitions of accuracy is not a measure. And I’d have more confidence in a decision made by a Koala. At least it’s cute!

–> Assign It

If we can’t even measure what it does, do you think we can measure how well the vendor pushing it really understands it? Definitely NOT!

And that, in a nutshell, is why there’s a lot of hot air and nothing of actual substance in all the AI discussions. Which you should avoid anyway. Because you don’t want tech with a 6% success rate [McKinsey, MIT], which is half the general success rate of new tech installations (now that tech failure rates have reached an all time of 88% [Bain]).

You cannot export-control math!

A truly brilliant observation by Mr. Stephen Klein in a recent post on how China May Be The Only One With Mythos because they may have saved its responses (and thus figured out how to replicate it).

(Gen-) AI LLMs are just mega math models. Really big mega math models with probabilities being computed on top of probabilities being computed on top of probabilities in force-feedback loops that reinforce its computations (which could be brilliant deductions or hallucinations that equal the acid high of the most LSD addicted junkie on the planet).

And since the outputs are dependent on the equations that define the model and the training data, if someone can recreate the training data they can reverse engineer the equations from the output if they are sufficiently adept at mathematics, or at least a close proximity.

Which means that blocking off access to those who can evaluate and improve the model will not help if those you don’t want to have the model already have it.

But this isn’t a post about the blocking of AI models (because I personally think that’s great), but a post about what happens if you try to hide your capabilities behind “proprietary algorithms based on math” assuming that no on else can recreate it if you don’t talk about it and that the IP alone justifies an unreasonable price for your product or valuation for your company.

Every country has mathematical geniuses, and more than one can come up with the next iteration of a mathematical theory at about the same time. Maybe only one gets remembered (Newton vs. Leibniz), but it doesn’t mean they didn’t both invent the same concepts at about the same time (in calculus).

And the more you trump it up, the more it entices someone to tear it down. (And figure out how you built it.)

On its own, the math alone is not the advantage you think it is. There are a lot of free scientific papers with great math. In fact, more than enough to create your own Claude, DeepSeek, Gemini, Grok, etc. But most people can’t because it’s not just the equations, it’s the parameters, the training data, and the implementation.

With regards to the implementation, just because you have server racks that can do trillions of calculations per second, that doesn’t mean you can code inefficiently. For example, an average token output by Claude requires billions of operations, and an average question posed to Claude will require 1,000 to 10,000 tokens to answer, for 1 trillion to 1 quadrillion calculations for an output. A poorly designed model could require 10, 100, or 10,000 times that.

The same goes for classical machine learning or optimization algorithms. Good implementations will require billions of calculations. Bad, trillions to quadrillions to quintillions. Responses go from real-time to hours to days to the computations never end.

Then there is the training data. You can’t judge a model implementation unless you have good training sets that are representative of real-world problems. Optimizing for theoretical problems that don’t exist in the real world isn’t helpful and may, in fact, lead to a worse solution that not even testing it at all!

The real differentiator is the expertise both in the implementation of solutions based on math and deep knowledge of the domain the solution is for. That can’t be recreated by mathematical ability alone and requires experience. And that can be export controlled (and represents the real value).

AI That Makes Recommendations Makes You Dumber, NOT Smarter!

Still too many posts about how great (Gen) AI is (in the age of LLMs).

They tell you truths:

AI can process all of your spend in seconds and find anomalies.

AI can collect all of the relevant market data and find opportunities.

AI can review large amounts of text and find risks or unfairly onerous contract clauses.

AI can track your SaaS utilization and ensure you are not being overcharged.

Yada Yada Yada.

All true, all fine.

But then the blasphemy starts. (Where I’m using the word in the context of the profane for the humanity bashing that it is.)

AI is great because it can recommend the spend “opportunities” you should pursue … and even automatically generate sourcing events for you.

AI can find the lowest cost when you need to spot buy and automatically buy/cut and send the PO for you.

AI can tell you what clauses to take out, what clauses to edit, and what clauses to add and automatically suggest the edits and write the new clauses for you.

AI can automatically enable and disable user accounts/seats, compute capability, application instances, etc. and save you money.

Now, theoretically, it can do all this. But practically, when it does so, it costs you money, capability, and you humanity.

When it selects opportunities, generates an event, and selects suppliers, it does so on a cost and historical utilization basis, with vacuum forecasts and whatever specs it can find. It doesn’t look at associated logistics costs, lead times, quality levels, service costs, certifications, safety, or anything else that is critical. You teach it cost, the metrics dictate that the CFO only cares about what shows up in the P&L, and you get the lowest cost piece of cr@p on the market. No big deal until customers get so fed up they start leaving, unless, of course it was the bolt holding the axels together on the bus that regularly drives the cliffs of the local mountain range or the door on the Jet you cram hundreds of passengers into.

When it selects the lowest price, there’s no guarantee the product will arrive on time or meet all of your requirements (because you just specified the cheapest card stock, but didn’t specify white and got pretty pink; 32 GB DDR chips, but didn’t specify they were for laptop upgrades and got server RAM; specified DBA, but didn’t specify Oracle and got someone who’s only ever used SQL Server).

When you tell it to slash your SaaS and Cloud costs, it happily deletes all the C-Suite accounts because they only log in once a month, cancels your vulnerability scanning service, because that costs way too much for something that happens only monthly, and deletes your main production database (because it cost way more than the QA database). And yes, plenty of news stories where it has already done all this.

When it scans the 60 page contract behemoth from the supplier, it overlooks the clause that transfers all liability to you for their AI failures (because you deployed the product) because you trained it on contracts where you transfer all liability of use of the equipment you create to your customers (because they use the product and accept not to use it beyond your specifications). Then when the AI accuses your best supplier of submitting fraudulent invoices, automatically files a report with the bank, which in turn freezes the suppliers accounts, which blocks all automated payments, which results in their energy supply being turned off for non-payment, which brings down their production line, which costs the supplier 2 million dollars, which results in them suing you … guess who’s on the hook when the court says “you can’t say the AI is responsible”?

But that’s just the direct costs.

The indirect cost is that it’s making you stupid.

The problem is this. Most of the time,

  • the opportunities will be real, not the most significant, but real, and the recommendations will be “good enough” that an average human won’t feel it worth the effort to qualify and/or improve
  • 80% of tail is usually well-defined cookie-cutter finished products/basket services and there will be enough description in the catalog/e-Pro system for the AI to get it “good enough” that the org can make it work
  • the security settings and “manual overrides” will usually prevent exec accounts and production instances from being deleted, and its ITs job to deal with the odd glitch, so who really cares
  • the missed risk won’t materialize in 95% of contracts, and usually not in the first 6 months

So people quickly stop questioning, start trusting, and then start blindly depending on it. The systems are allowed to do whatever they want as long as they aren’t creating fires worse than what the humans are already dealing with. And even as performance degrades, requirements change, or system costs escalates, nothing is checked, and as slightly worse decisions are constantly reinforced and bad data piles up, the organization is sitting on a ticking time bomb. Which is going to go off. The only question is how much damage it is going to do.

Which you won’t be able to deal with when it does go off because you won’t have a clue what to do.

Every time you fail to question it, your cognitive skills atrophy.

Every time you fail to use a skill, your expertise dissipates.

Every time you fail to put the effort in, your problem solving stamina degrades.

Every time you fail to think about the situation, your knowledge erodes and you forget.

When they say your mind is a muscle, they’re not exaggerating. Bodybuilders don’t work out every day just to build, they do it because it’s a basic requirement to maintain what they have so their muscles don’t atrophy and their hard work dissipate.

The brain works the same way, when you don’t use it, it atrophies.

Numerous studies have shown what happens when you use AI even for the simplest of tasks. Even typing vs hand writing reduces brain power. Using it to edit reduces more. Using it to write reduces more. (15% or more — you’re effectively shaving 15 points off your IQ!) Using it for strategy gets to the point where you might as well be boarding the short bus and taking the remedial class, because in a matter of months you’ll have trouble keeping up with that!

So, the more AI you use, the dumber you get.

And that’s great?

Maybe if you’re a malevolent billionaire trying to dumb down humanity to the point that they are too stupid to question your true intentions, but otherwise …

If You’re Spending 250K Annually Per Engineer On AI …

Then not only are you contributing to planetary destruction (through the generation of between 1.32 tons (high end models, 1 joule per token) and 84 tons (low end models, 2 joules per token) of CO2 to power those data centres, which is about 0.2 to 12.7 times the average individual carbon footprint, with an expectation of 7 to 11 tons (Source), and the utilization of 300,000 gallons to 5,000,000 gallons of water a day to keep those servers cool, or a town’s worth of water every day!

BUT YOU ARE NEEDLESSLY WASTING 400K+ A YEAR

1. Less than 20% of AI generated code survives unscathed in a commercial enterprise software product once senior developers weed out all the security errors, boundary condition errors, and generated code that doesn’t even solve the problem. So, that’s 200K of 250K down the drain as only 20% of output is usable.

2. Having to fix AI generated slop will consume 80% of a good senior developer’s time — a developer you should also be paying 250K a year.

End result, you’ll losing 200K + 200K per developer you force AI coding tools upon!

But hey, it’s your money. If you want to p!ss it away so NVIDEA’s CEO can get richer selling more CPUs we don’t need, that’ up to you!

The linked article contains some metrics, but here are a few others.

  • token prices vary widely, from an average of around 50c/M tokens on the smallest, cheaper models to $75/M tokens (or higher) for higher end “workhorse” models
  • energy processing requirements per token are estimated to be between 1 joule and 2 joules
  • you can buy 14.3 Trillion tokens at the median of around $17.5/M tokens (and 35 times that at the lower end)
  • processing 14.3 T tokens will take about 4000 kwH @ 1 joule/token
  • on an average NA grid, expect to produce 500 to 600 g of Co2 per kWh (since most of our grids are still dirty)

The Bullshit Filter for Enterprise AI Startups consists of 12 Questions!

Not 11!

Backing up, earlier this year Jason Busch published his 11-Question Bullshit Filter for Enterprise AI startups. This was, and is, needed because the vast majority of Enterprise AI startups are bullshit (especially in FinTech and Procurement) and the sooner you figure that out, the better.

I was hoping that, by now, the AI startup scene would start crashing due to over investment, lack of returns (only 6% of AI implementations have generated an ROI), and, generally, lack of usefulness. (AI can serve up your data, show you complexity and even help with automating some tasks, but it can’t make decisions and, due to lack of anything close to intelligence, can’t even do basic tasks without your oversight.) But, even worse, these solutions are still multiplying like Fibonacci’s rabbits and their claims are getting more outlandish by the day. (How many times do we have to tell you AI Employees Aren’t Real, you should NOT engage any vendor selling “AI Employees”, because you definitely do NOT want AI Employees.)

So, since they are flooding our space with BS marketing and making ridiculous claims about what their useless apps can do, it’s more critical than ever that you be able to suss out the BS claims from the non-BS claims. (Hint: 95% are BS claims, so it wont’ be easy!)

We’ll start with Jason’s 11 filters, which we’ll number 12 down to 2, because he left out the most important filter, and the one that, if it fails, allows you to skip the next 11.

Filter 12: Founder DNA
Can they build and sell? Likely not. Chances are, if they’ve cut through the noise and reached you, they can only sell. And if you did find a builder, they won’t survive long enough to support you if they can’t sell.

Filter 11: Motivation
Is failure unacceptable? (Every startup team will say it is, but unless every founder has a reason they simply cannot accept failure, when the going gets tough … the tough get going … and quit.)

Filter 10: Interface
Is it designed for those who will ACTUALLY be using it?

Filter 09: Categorization
Does the product actually do something new? Is there a strong reason for the market to adopt it?

Filter 08: “Found Money”
Are there instant benefits that sell themselves on the first demo.

Filter 07: Displacement
Does the product workaround or replace a solution that buyers hate?

Filter 06: Functional Bonds
Does the solution cross boundaries that increase value beyond peers?

Filter 05: Data Delta
Is there a “data” strategy to exploit the delta between what humans can easily consume and what AI can leverage (and summarize into something useful for human data ingestion)?

Filter 04: “Messy Middle”
Can the solution ingest external “dark data” and turn it into actionable insights without requiring a(n extensive) manual data-cleansing project? (Quick review and correction is okay.)

Filter 03: Connect the Dots
Does the app bridge the gap between “Watercooler Data” and “System of Record Data” (ERP/PO) to explain the why behind an analysis or recommendation?

Filter 02: “Show Your Work” Audit
Can the user drill into any output, see each and every step the AI took, drill down to the source data, and verify that everything is correct, accurate, and no data was changed?

These are all great filters, but there’s no point going through them if you don’t check the most important filter first:

Filter 01: Is it LLM-based?
If yes, move along. Don’t waste any time.

Most of the failures in the age of AI come from Gen-AI LLMs that promise the world and don’t even deliver a pile of dirt. That hallucinate on every other query. That burn up thousands of dollars of tokens to deliver less than fresh MBA interns with no real world experience and no clue to share on their first day no less.

Even worse, the majority of these players are simply wrapping third party LLMS in the creation of their “solution”. That’s not a solution at all. That’s an unmitigated disaster waiting to happen!

In the rare case an LLM actually offers a partial solution, it is best to go straight to one of the major providers. That way, you know who’s responsible when something goes wrong and don’t have to worry about providers playing the blame game and pointing fingers at each other.