thomas.wieberneit@aheadcrm.co.nz
Usage-Based Pricing for Copilot Is Good for Microsoft’s Investors. Read That Sentence Again.

Usage-Based Pricing for Copilot Is Good for Microsoft’s Investors. Read That Sentence Again.

TheStreet ran a piece this week arguing that, of Microsoft’s two Copilot announcements, the shift to usage-based pricing matters more to investors than the DeepSeek flirtation. That read is correct. It is also the tell. Here is what Microsoft actually did. Copilot Cowork, the agent that reaches across Microsoft 365 to run multi-step work on your data, is coming off the flat per-seat add-on and moving onto consumption billing the company calls “Copilot Credits.” Charles Lamanna, who runs Copilot, told Axios the product could not be offered on an unlimited-use basis. The users he pointed to are the ones doing hundreds of tasks a week. He called them “way productive.” And then he said the part vendors normally keep off the slide: their costs go very high. So the most productive users are the expensive ones. Hold that thought, because the whole argument lives there. What “good for investors” is really saying A pricing model earns the label “good for investors” when three things are true. Revenue starts to track cost-to-serve. Revenue scales with consumption instead of sitting flat per seat. And the vendor stops eating the margin on its heaviest users. All three are true here. None of them is a statement about whether a customer got value. That is the gap I want to sit in for a minute. Usage-based pricing meters an input. Tokens, compute, credits, whatever the unit. The customer does not buy tokens because they want tokens. They want a finished report, a resolved ticket, a reconciled spreadsheet. The token count is the cost of producing the outcome, not the outcome. And the relationship...
Pega’s fix for runaway AI costs: stop the agents from thinking at runtime

Pega’s fix for runaway AI costs: stop the agents from thinking at runtime

The news At its PegaWorld conference in Las Vegas on June 8, 2026, Pegasystems announced Pega Infinity 26, which it says will be available in Q3 2026. The principal change is commercial: Pega is moving away from per-token pricing for its AI agents toward a flat charge per completed “case,” which it defines as a task carried out from start to finish, such as a customer changing an order, a loan approval, or a claim. Pega frames the move as removing what it calls the “AI token tax“. The pricing change rests on an architecture Pega calls Predictable AI. Reasoning-heavy AI work is concentrated at design time, when workflows are authored in Pega Blueprint and the new Infinity Studio. At runtime, a lighter-weight model identifies the user’s intent, selects a pre-approved workflow, and executes it step by step; where an individual step requires a language model, for example to parse a document or summarize a prior interaction, that step is given bounded instructions rather than open-ended latitude. Pega gives two reasons: more consistent outcomes, because agents follow approved workflows rather than re-reasoning each request, and more predictable cost, because the heavier processing happens only once during design rather than on every transaction. The architecture is not new to this release. Pega introduced Predictable AI Agents in May 2025 and integrated them into Pega Infinity ’25, which reached general availability in December 2025. Infinity 26 primarily adds the outcomes-based pricing model, alongside a companion announcement that exposes Pega processes as Model Context Protocol (MCP) servers, allowing third-party agents from Anthropic, OpenAI, Google, and AWS to call them under Pega’s governance...
The Illusion of Value: Why Salesforce’s Agentic Work Unit is the New “Bad Query” of the AI Era

The Illusion of Value: Why Salesforce’s Agentic Work Unit is the New “Bad Query” of the AI Era

The News On February. 25, 2026, Salesforce announced a pricing and metrics update. During the company’s Q4 FY2026 earnings call, CEO Marc Benioff, together with CMO Patrick Stokes, unveiled the Agentic Work Unit (AWU). Positioned as a metric to quantify the labor performed by autonomous digital systems, Salesforce defines an AWU as one discrete task accomplished by an AI agent. According to Salesforce, this discrete task represents the exact moment “raw intelligence is converted into real work“. It is not a fixed unit but measured as a processed prompt, a completed reasoning chain, or an invoked tool. Salesforce explicitly designed the AWU to move the industry conversation away from the raw consumption of Large Language Model (LLM) tokens. As Benioff noted, tokens only measure “how much an AI talks,” whereas the AWU is intended to measure actual business execution. The scale of this rollout is massive. Salesforce reported that its platform has already processed over 19 trillion AI tokens, translating them into 2.4 billion Agentic Work Units, with 771 million AWUs delivered in the fourth quarter alone. This new metric serves as the underlying foundation for Salesforce’s evolving Agentforce monetization strategy. The bigger picture Following a nearly 18-month period of pricing triangulation, which included a $2.00 per conversation model and a $0.10 per action “Flex Credit” model, Salesforce is leveraging the AWU to track system utilization, even as it wraps enterprise purchasing in familiar, unmetered per-user license agreements starting at $125 per user per month.   To understand the significance of the Agentic Work Unit, one must view it through the lens of a broader industry crisis: the so-called “SaaSpocalypse”...