Every wave of customer-facing technology arrives with the same promise: this time the balance of power shifts toward the customer. As we discussed in our recent CRMKonvo with Dan Miller, founder and analyst emeritus of Opus Research, voice AI is making that promise again, and this time both sides of the counter are armed. The more interesting questions are where the customer's weapon is kept, who pays its rent, and who gets to whisper to it while it works.

TL;DR

If you want to watch the full CRMKonvo, please go ahead here (optimized for smartphones) or here (optimized for tablets/computers).

Else, be my guest and continue to read.

Or do both …

Automated Voice Was Never Built for You

Dan has watched this market almost from its beginning, and he is blunt about its origin. Automated voice entered the contact center "almost entirely for cost savings". Everything after that was repair work: better speech recognition, task-specific tuning for finance and travel, and a long campaign to make a system optimized for cost do a passable impression of service.

The same instinct shows up in voice analytics. Once recordings could be mined, the first trigger enterprises asked for was a competitor's name, then a churn score. The retention arithmetic was well known; it’s roughly six times cheaper to keep a customer than to win one, and the wisdom that followed "got distilled into a few rules". Detect the person about to leave, offer them something.

In blunt words: the industry learned to close the door, not to fix the room. Nothing about generative AI repeals that. It runs the same rules faster, on more data, with better grammar.

The Same Data, Two Completely Different Businesses

Proverbially, technology is neutral and the intent behind the listening is not. Reading and analyzing a customer's conversations to make the customer the means to get revenue and reading them to make the customer successful with revenue as the consequence, are two mindsets that produce very different decisions from identical data.

Dan looks for the win-win and argues the crossover point is real: "successful firms do a better job of not exploiting their customers", and you cannot save your way into profitability. In a contestable market he is right. The trouble is how many markets are not. Starlink raises prices for the stated reason that it can, retires a plan and sells the stranded customers a dearer one. Banks make you dependent first and negotiate second; try collecting your wages in a brown envelope this year. Volkswagen shipped an ID.7 with a single set of rear window switches and a toggle to choose which window you are operating, because somebody found a few cents in the parts list.

None of that is a data problem, and no conversational layer fixes it. There are two ways to make a profit: by making me successful, or by making sure I make you successful. They are not mutually exclusive, but the slider between them is set by the executive, not by the model.

Your Personal Agent Has a Landlord

Dan's constructive answer is quite interesting. Build your own agent: one that knows your preferences, your payment methods, your preferred vendors, and goes into the market with your horsepower rather than the vendor's. Assume everyone else is a wolf. "I want this thing that is all mine and I can do battle on a better than even playing field."

He named both problems himself. The first is assembly. The people extracting real value from these tools are already technologists; everyone else gets a chat window. Do I buy the Mac with Apple silicon so I can run my own model? We have had the conceptual framework for a voice agent on a smartphone since Siri in 2009: seventeen years of the same slide, still no delivery.

The second problem is structural. Draft regulation in the United States already assumes a trusted, custodial user agent hosted on somebody's infrastructure: Google, Apple, your card issuer. Dan sees the consequence clearly: "there would be room for my agent to be influenced by these entities that don't have my best interest at heart." Preferential placement, promotional dollars, ad support, and then the slow enshittification Cory Doctorow named.

This is not speculation, it is a rerun. Social media was sold on exactly this rebalancing. The consumer held control right up until vendors understood what the tools were worth and bought the position back, because they had the funds. An agent you do not host is an agent you do not own.

If Agents Talk to Agents, Voice Shrinks to the Handoff

Here is my question for the voice AI market. If agent-to-agent commerce becomes the mainstream, why would two digital agents converse in a human language at all? Voice then survives only at the two edges: where I express intent, and where the result is handed back to me.

Dan broadly agrees. Ninety percent can be machine to machine and invisible; the remainder is rendered as speech, because it is the most natural interface for a human. Which is awkward, because the edges are exactly where voice has always broken. Ralf raised the automatic barista a Swiss food company built eight years ago: superb coffee, a microphone and speaker on the front, no chance of surviving a vandalised station concourse. Demo quality is not operational quality. Dialects, background noise and the oldest failure of all, which Dan still states best: "The last thing you want is going around and around because it didn't understand you."

The evidence on convenience is decades old and nobody likes it. Amtrak ran one of the first automated phone booking systems, and users programmed the touch tones into their speed dials. The fastest route through the voice system was not talking. Meanwhile customers consistently choose to sit on hold for a human rather than deal with the automation. That is a revealed preference, measured over thirty years, and it has barely moved. The reason is simple: The systems are designed about the company’s wants, and not the customers’.

Dan's engineering answer is the unglamorous one: list the ways the thing will fail, find the three to six modes that cause ninety percent of the failures, and fix those. It is the right answer. It is also not what the current round of voice demos is tuned for.

Intent Fulfilment Is the Only Meaningful Metric That Can Be Used By the Vendor

Dan's ice maker stopped working. He photographed the model plate, asked an LLM, got a wrong answer, sent more pictures, and eventually landed on unplugging the unit for thirty seconds. It worked. No ticket, no 800 number, no technician, no survey. Service happened and the manufacturer has no record that it did. Existing service delivery infrastructure, in Dan's words the people, the systems and the documentation, is what LLMs will eat.

Which is why the argument in Mitch Lieberman's forthcoming book, The Conversation is the Record, matters more than the agent hype around it. Put the conversation, with the context wrapped around it, into the system of record across every channel, and you can finally ask the one question worth asking. Dan's lightning-round metric was exactly that: intent fulfilment. Commitment made, commitment kept.

One can argue that a fulfilled commitment is not a finished job, because commitment doesn’t necessarily cover the intent. It still beats everything currently on the dashboard. Handle time, containment, deflection rate and NPS all measure how the enterprise felt about the interaction. Intent fulfilment measures whether the customer got what they came for, and it produces an auditable list of promises the company did not keep. That is precisely why no vendor is rushing to sell you the scoreboard.

Pragmatic Playbook for Enterprise CX Buyers

Asked who gains more influence over the next three years, Dan was honest: in his world, customers; in the real world, a standoff. Standoffs are decided by preparation.

For enterprise technology leaders evaluating Voice AI and agentic platforms, the market is currently a minefield of over-promised capabilities and hidden architectural debt. To avoid making a costly mistake, buyers must ground their strategy in three core operational rules:

Demand Intent Fulfillment Metrics, Not Deflection Rates. Stop measuring your CX success by how many calls your Voice AI prevents from reaching your contact center. Deflection is a cost-center metric that frequently masks customer frustration. Instead, instrument your systems to track commitment resolution: did the AI correctly capture the user's intent, trigger the appropriate backend API, and resolve the issue end-to-end? If your AI cannot execute real-world workflows inside your ERP and CRM systems, do not deploy it to customer-facing channels.

Decouple Conversational Data from Proprietary Vendors. Do not allow your Voice AI or contact-center-as-a-service (CCaaS) vendor to lock your interaction data inside their black box. Implement vendor-neutral conversational data standards such as vCons. Containerizing your audio, text, and metadata ensures that you retain full ownership of your customer interaction history. This allows you to audit AI performance objectively, feed pristine context into your RAG pipelines, and switch underlying LLM providers as better models emerge without losing your historical memory.

Establish Frictionless Human Escalation with Full Context. AI agents must never become digital dead ends. Design your conversational architecture so that the moment an agent detects intent ambiguity, user frustration, or an out-of-bounds request, the interaction transfers immediately to a human agent. Crucially, the human representative must receive the complete (Vcon container) transcript and real-time intent summary instantly, eliminating the infuriating customer experience of having to repeat information to a human that was already given to a bot.

Listening is cheap. Keeping your promise is the product.