Every hype cycle produces its own comfortable silence, and this one has a good one: almost everybody is doing AI, and almost nobody will say out loud whether it worked. In our recent CRMKonvo with Jon Reed, co-founder of diginomica, we spent an hour poking at that silence, with Ralf Korb doing the poking alongside me. Jon is one of the analysts who actually tests the thing before writing about it, which makes him tiresome company for vendors and excellent company for buyers. The conversation did not land on whether AI works. It landed somewhere considerably more uncomfortable: most enterprises cannot say what working would look like, and they started spending anyway.

TL;DR

If you want to watch the full CRMKonvo, please go ahead here (optimized for smartphones) or here (optimized for tablets/computers).

Else, be my guest and continue to read.

Or do both …

Sixty Percent, And Nobody Is Blushing

Jon opened with numbers rather than opinion, a habit more of us should copy. McKinsey's recent work on AI measurement found that nearly eight in ten companies are using generative AI in some capacity, while around sixty percent report not seeing enterprise-wide EBIT impact from those programs. The gap between activity and impact is apparently not closing. Instead, it seems to be widening. A small group of over-performers is automating end-to-end workflows inside specific domains and getting results, and even they argue about what to measure and how to attribute the improvement.

Let that sink in for a moment. These are organizations that committed budget, headcount and executive credibility to a program, then discovered they had never agreed on a definition of success. Jon asked the obvious question: "Why would you undertake a project like this if you had no idea how you were going to measure the success of it?"

The answer is not stupidity.

It is fear.

There is "a profound fear of missing out or being left behind", and the people applying that pressure are usually the ones furthest from the technology. Executives and board members "have some of the most unrealistic ideas about AI in the entire organization". So the CIO is told to spend on something, anything, and the something arrives in the shape of a forward deployed engineer. Fine role, wrong instinct: a forward deployed engineer is a technologist. They do not know your business, and this was never an engineering exercise.

The Model Stopped Being the Product

This is the part of the story most vendor keynotes skip. Five or six years ago, scaling language models produced results that looked like emergent intelligence, and the valuations followed: if you can build truly cognitive systems, you can replace large parts of the workforce and justify almost any capital expenditure. What actually arrived is "a facsimile of intelligence". Then the scaling laws slowed, the training data ran thin, and investors got nervous.

What the labs did affects you more than the models themselves do. They wrapped the models in compound architectures with external verification, symbolic tools, deterministic cross-checks that refuse to let an agent post to the general ledger when the entry fails a rule. The industry sells this as context engineering and harness engineering. Strip the vocabulary away and the model has become the least interesting component in the stack. The architecture around it has become the product.

Jon's summary of the exercise is simple: "We're taking a tool that was not intended to be deterministic, and we're trying to see how far we can push that." That is a candid description of the state of the art, and it carries a warning no vendor slide will show you. Guardrails work in one direction only. "It's a lot easier to stop agents from doing something wrong than to know for sure that they did something right, because they don't understand what the right thing is." Blocking a bad output, better one too many than missing one, is engineering. Certifying a good one is still your problem.

Expertise Is Not a Commodity, and Nothing Is Learning

Two claims circulating on LinkedIn got taken apart, and both had it coming.

The first is that expertise has been commoditized.

It. Has. Not.

What has been commoditized is working-level knowledge across many domains, which the models absorbed during training. Working knowledge is not expertise. "It's only expertise that can identify the problems in the model output," Jon argued. That sentence should reorganize your hiring plans. If you believe the machine is the expert because it passed the bar exam, you will ship its mistakes at scale and file the result under productivity. Mathematics is an exception, because synthetic data works inside the closed confines of maths. Your industry is not maths.

The second claim is that the agent learns from your users.

It does not.

The language model is not learning while you talk to it. It is pre-trained, then trained, then frozen, and adjusting weights on the fly runs into catastrophic forgetting; Rich Sutton has a Turing Award and a working explanation of why. When a vendor says the system learns continuously, they mean that a knowledge graph or memory store are updated with your preferences. That is a storage mechanism disguised as learning. It is dangerous because it gives buyers the wrong idea of what is possible, and having wrong ideas about the possible is how budgets get burned.

The CSAT Trap: When Nothing Got Worse Counts as a Win

Now to CX, where the reasoning gets worse rather than better. We looked at companies that replaced level one service with AI assistants, reduced headcount, and reported the result as a win because their CSAT scores did not go down. Hold that up to the light. The stated ambition was to change nothing about how customers experience you while spending less on them.

"Fine, but you're not Amazon." Unless you are a behemoth or an airline, service is one of the few places where you can still out-compete companies that have more money than you. The question is not whether the bot held the line at nine in the evening. It is whether these tools let you run the best service in your industry, including at nine in the evening when your people have gone home.

Corporate intent decides that outcome. If the corporate desire is to solve issues, the technology gets designed to solve issues. If the desire is to deflect them, you have bought a deflection machine with better grammar. And the failure mode is almost never the answer itself; it is the escalation. Jon's own pharmacy routes him through a voice system that offers to help, asks him to describe the problem, then loops him back into the automation he was trying to escape: "If the automated system had answered my question, I wouldn't be asking to talk to the frigging pharmacist." Every enterprise reading this has built that loop somewhere.

Not every customer warrants the same treatment either, and pretending otherwise is not fairness, it is laziness. Your largest account should probably not be routed into the same voice system as everybody else.

Architecture Follows Intent

The best line of the hour was not Jon's own. He borrowed it from a diginomica colleague writing about the International Rescue Committee's AI operating model: architecture follows intent. Decide who you want to be, then build the thing that makes it possible. Most enterprises run that sequence backwards, buying architecture and hoping an intent turns up later.

Equifax came up as the counter-example, and the detail is the useful part. They credit their AI results not to a clever agent but to five years of cloud migration that left their data in a standard fabric, plus proprietary data the models have never seen. Nobody sensible will tell you to spend two years modernizing before touching AI. Jon did not, and neither will I. But the modernization track and the AI track run in parallel, and the sprinkle-sauce theory, the one where AI lets you skip the discipline, is "a LinkedIn feed fantasy land".

Pragmatic Playbook for Enterprise CX Buyers

Settle three things before the next AI proposal reaches your desk.

Build the evaluation suite before the program office. You have to have transparency over what your AI is doing. Define the business outcome, the baseline and the attribution method before the contract is signed, not after the pilot disappoints. Pick a problem meaningful enough to matter and contained enough that getting it wrong does not break the business. If nobody in the room can state success as a number, you are not ready to buy.

Put your pricing and your data in the contract. Any change to outcome-based or consumption-based pricing requires six months of notice so you can adjust. Moving off user-based licences to pay for tokens with no business result attached is not an advancement, it is a higher invoice. And when the vendor says their agent learns from your users, ask these two questions: how exactly does it learn from my users, and how do you protect that data? An update to the knowledge graph is not learning.

Design the escalation first, then hire someone to check the whole thing. Most customer anger at AI support is not about the answer, it is about being unable to get out. Build the route to a human before you build the deflection, and keep your most valuable accounts out of the automation entirely. Then consider the role Jon would add to the org chart: an AI ombudsperson whose job is to walk into departments, gut-check what is being built, and flag the vulnerabilities and the opportunities nobody else is positioned to see.

Architecture follows intent. Buy the intent first.