NextAI+ Praxis--:----UTC
Enterprise AI Deployment Signals

Is More AI Usage Always Better? Enterprises Begin to Question the Real Return of Agentic Workflows

28 July 2026
Long read · 14 min
By NextAI+ Praxis

Recently, when enterprises measure the progress of their own AI deployment, they often first look at several types of indicators: how many employees have started using it, whether model call volume is rising, whether token consumption is growing, and whether more agentic workflows are emerging internally. In the early stage of AI entering enterprises, these numbers can indeed explain something, namely that AI has moved from experiments by a small number of people into broader work scenarios. But after AI begins to connect to knowledge bases, read business context, call system tools and empower specific processes, usage itself also starts to become a new cost problem. Enterprises are gradually realizing that the fact that AI can complete a task only proves that it is technically feasible; whether this task is worth being completed by AI still needs to be measured again against cost, efficiency and business outcomes.

A group of recent signals is also bringing this problem to the front. Databricks has launched AI spend controls to help enterprises limit runaway AI spending; OpenAI has begun to provide more detailed usage analytics for ChatGPT Enterprise (a higher-level commercial AI offering provided by OpenAI); platforms such as Microsoft and Salesforce are also remeasuring enterprise AI consumption through credits, actions and tool calls. Related research shows that token consumption in multi-step agent tasks inside enterprises is far higher than that of ordinary Q&A, and higher consumption does not necessarily bring higher accuracy. Taken together, these signals show that enterprise AI deployment is facing a more realistic question—under technically feasible conditions, can the consumption generated by the continuous operation of AI really bring sufficiently clear business returns?

On this cost-benefit issue of AI deployment, this article has two core judgments:

  1. The evaluation criteria for enterprise AI deployment are shifting from usage scale to workflow return. Employee usage rate, model call volume and token consumption can show that AI is being adopted, but they cannot by themselves prove that AI is creating business value. As agents enter real processes, enterprises need to measure AI consumption, labor savings, process cycle time, output quality and risk changes within the same framework.
  2. The key to agent cost governance lies in redividing the task boundaries among models, automation and human labor. Enterprises need to decide, according to task complexity, frequency, risk and business value, which links are suitable for stronger models, and which links are suitable for lower-cost models, rule-based automation or human judgment. Without workflow-level cost attribution and outcome measurement, usage growth may instead obscure the real return of deployment.
§ i

Vendors Begin to Push AI Costs Down from Account Subscriptions to Agent Actions

The starting point of the enterprise AI cost problem is first pushed to the market by software and platform vendors. Traditional enterprise software usually uses accounts, seats or subscriptions as the main pricing unit, and what enterprises purchase is the right for employees to enter a system and use a certain software function. This logic fits the era in which people use software: costs mainly revolve around “how many people can use it,” while how many times a specific employee clicks in the system or how many actions they complete usually does not become an independent pricing object.

After agents enter workflows, vendors have begun to break pricing units into finer pieces. Microsoft mentioned in Announcing the New Work IQ APIs, published on its official website, that after Microsoft Work IQ API became generally available on June 16, 2026, it adopted a consumption-based pricing model based on Copilot credits, without a separate subscription, SKU or per-user license. When enterprises call Work IQ API to access work context and application capabilities in Microsoft 365, they will directly consume Copilot credits.

Microsoft 365 admin center Cost management page for Copilot, showing spending policies, credit limits, and a Work IQ API policy row with its monthly spending limit and credits available.
Figure 1: The Copilot Credits cost management interface in the Microsoft 365 admin center. Source: Microsoft, compiled by NextAI+ Praxis.

The flexible credits of Salesforce’s agent platform, provided by the American customer relationship management company Salesforce, point to a similar change. Salesforce’s official help page states that its agent platform adopts a “pay per action” flexible credit model, in which each agent platform action consumes 20 flexible credits, about $0.10; these actions can include updating customer records, automating workflows or resolving cases. Salesforce’s pricing page also gives an example: if a use case contains 3 actions, it will consume 60 flexible credits; when 100 users each handle 3 cases per day for 20 days per month, the monthly example cost is $1,800. This pricing method turns the business actions executed by agents into measurable units, and also allows enterprises to begin estimating AI costs according to task frequency, number of actions and process design.

Salesforce Agentforce pricing page comparing three tiers: Salesforce Foundations at $0, Flex Credits at $500 per 100k credits billed per action, and Conversations at $2 per conversation.
Figure 2: Salesforce Agentforce measures agent usage through Flex Credits and Conversations. Source: Salesforce, compiled by NextAI+ Praxis.

The changes from Microsoft and Salesforce show that enterprise AI costs are expanding from relatively fixed software entry points into a hybrid structure of “basic usage rights + consumption credits + agent actions.” Vendors first break pricing units into finer pieces, and enterprises then have to reinterpret AI usage: whether an agent is enabled is only the starting point of cost generation; how much context it reads, how many tools it triggers, and how many actions it executes will determine the continuous cost of this workflow in real operation.

§ ii

As Agentic Workflows Lengthen, AI Spending Pressure Begins to Emerge

After vendors break AI costs into credits, actions and context calls, enterprises will soon encounter the next layer of problem: agent tasks themselves are more likely than ordinary Q&A to amplify consumption. Ordinary Q&A usually revolves around one input and one output, and the cost is relatively easy to estimate; after agents enter business processes, tasks are stretched into a continuous chain, including reading context, planning steps, calling tools, waiting for returns, reasoning again, verifying results and retrying after failure. Each link generates new input, reasoning and verification costs, especially in long-context scenarios such as coding, legal, customer service, sales and procurement, where the consumption of a single task may be far higher than enterprises’ initial understanding of “using AI once.”

Related research has already provided a more concrete explanation. In April 2026, experts and scholars from the University of Michigan, Stanford University, MIT, Google DeepMind, Microsoft AI and All Hands AI jointly released a research paper on AI spending, How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks. This study analyzed the execution trajectories of 8 frontier large models when performing agentic coding tasks, focusing on where token consumption occurs in the task chain. The study found that token consumption in agentic coding tasks is about 1,000 times that of ordinary code reasoning and code chat, and the token consumption of the same task can vary by up to 30 times across different runs; higher token usage also does not stably bring higher accuracy, and accuracy often peaks in a medium-cost range before tending to saturate. Although this study focuses on the token usage scenario of coding agents, it reveals a transferable mechanism: the closer agents get to real task execution, the more easily costs spread along the task chain, and the harder it becomes for enterprises to accurately estimate final consumption before a task begins.

In response, the platform side has begun to add direct spend monitoring and control capabilities. In June 2026, in an exclusive report by the digital media company Axios, the American data and AI company Databricks announced the launch of an AI spend control tool called “Unity AI Gateway,” emphasizing that it can set budget limits, identify abnormal consumption, and manage usage costs across model vendors. The starting point for launching this tool came from the “painful lessons” of its enterprise customers in AI spending—there had been cases where tens of millions of dollars in AI spending were accidentally generated within a single month. See Exclusive: Databricks Rolls Out AI Spend Controls.

Axios technology report stating that Databricks is launching tools to help companies cap AI costs after customers accidentally spent tens of millions of dollars on their AI bills in a single month.
Figure 3: Axios reports that Databricks launched AI spend control tools. Source: Axios, compiled by NextAI+ Praxis.

At the same time, OpenAI is also adjusting in the same direction. In June 2026, OpenAI launched more refined usage analytics and spend control functions for ChatGPT Enterprise, allowing enterprises to view credit consumption by different users, products and models, and to set corresponding credit management. See OpenAI Introduces Enhanced Usage Analytics, AI Spending Controls for ChatGPT Enterprise.

Bulleted summary of OpenAI's new ChatGPT Enterprise admin console features: consumption breakdowns by user, product and model, usage trend monitoring, default workspace credit limits, per-group limits, and employee credit requests.
Figure 4: OpenAI launched usage analytics and spend control functions for ChatGPT Enterprise. Source: OpenAI / Reuters, compiled by NextAI+ Praxis.

Correspondingly, the enterprise side has also begun to take restrictive actions. A Financial Times report in June 2026 mentioned that companies such as Amazon, Walmart, Cisco, Uber and Meta have begun restricting employees’ use of AI tools to cope with rapidly rising AI costs; among them, Uber has already set a token limit of $1,500 per employee per month (previously, the company had used up its entire 2026 AI budget by April). Walmart also restricted the use of its internal AI coding tool “Code Puppy,” one reason being that employees were making large numbers of similar requests for repeated questions, causing cost and resource waste. Enterprises still hope AI can improve efficiency, but when large numbers of employees and agents continuously call models, management will start requiring usage behavior to enter budget, quota and permission management systems.

Taken together, these signals show that AI spending pressure does not come from a single model call itself, but from the continuous consumption after agentic workflows are lengthened. In response, vendors and enterprises each have their own adjustment methods: the former begin to provide functional services for spend control and usage analytics; the latter begin to set quotas and limit repeated calls to cope with the spending pressure continuously generated by AI operation.

§ iii

Enterprises Need to Redivide the Task Boundaries Among Models, Automation and Human Labor

Spend control and usage analytics can help enterprises see where AI costs come from, but they cannot automatically tell enterprises how to redesign workflows. The spend monitoring provided by Databricks and OpenAI, and the token limits and quota management set internally by enterprises, solve the question of “whether consumption is visible, whether it exceeds budget, and whether it needs to be restricted.” But when enterprises truly want AI to continuously enter production processes, the next step is not simply to reduce call volume, but to return to the task itself: why certain tasks continuously consume high-cost models, whether these tasks should continue to be executed in the same way, and whether they have more suitable ways to be handled.

Therefore, the object of AI cost governance should not stop at the total bill or a single call, but should move down into the task nodes inside the workflow. A seemingly complete AI process is usually composed of multiple links with different properties. Taking customer service as an example, the same process may include issue classification, historical record retrieval, policy matching, response generation, abnormality judgment and human escalation; taking sales support as an example, AI may need to organize customer background, generate communication materials, query product information, match pricing rules and write back to CRM. The complexity, frequency, risk and business value of each node are all different; if an enterprise defaults to using the same model to handle the entire process, costs can easily be amplified in certain low-value or high-frequency nodes.

Therefore, from a consulting perspective, this article can roughly divide enterprise workflow tasks into four categories from the dimensions of task value, complexity, frequency, rule stability and responsibility boundary:

  1. High-value, high-complexity and high-risk tasks can bear higher model costs. Complex legal review, key customer proposals, code refactoring, compliance risk analysis and complex financial judgment all require professionals to invest a great deal of time, and the cost of error is also higher. In such tasks, the value of stronger models is not only to generate a draft, but to help experts organize materials faster, identify risks, compare options and form a basis for judgment. As long as what it saves is high-value expert time, or what it reduces is uncertainty in high-risk links, higher model consumption may become a reasonable investment.
  2. High-frequency, standardized and repeatedly occurring tasks need to be reexamined even more. Meeting minutes, ordinary document summaries, standard customer service replies, sales email drafts, simple contract clause extraction and ticket classification may not look costly in a single instance, but because they occur at high frequency, cumulative costs can quickly expand. If such tasks are long-term defaulted to the strongest models, enterprises can easily consume large amounts of budget in low-risk, low-differentiation links. A more sustainable approach is to introduce smaller models, fixed templates, caching, batch processing or structured retrieval, allowing these tasks to be completed at lower cost within an acceptable quality range.
  3. Tasks with stable rules and little room for judgment are more suitable to return to traditional automation and system interfaces. Order status queries, inventory availability queries, field validation, format conversion, fixed approval reminders and system notifications usually have clear inputs, stable rules and predictable outputs. Enterprises do not need large models to re-reason over these tasks every time. Database queries, API calls, rule engines, RPA or fixed workflows are often cheaper, more stable and easier to audit. AI can exist as an entry point or explanation layer, but the underlying execution is more suitable for deterministic systems.
  4. Tasks with major responsibility and ambiguous boundaries need to retain human leadership. Major customer negotiations, strategic supplier selection, highly sensitive legal opinions, complex organizational conflicts and management decisions involving major responsibility all require the combination of relationships, context, experience and accountability. AI can help organize materials, retrieve historical information, prompt risks and generate alternative options, but final judgment still needs to be completed by humans. For such tasks, excessive pursuit of continuous automation may create new unclear responsibilities and judgment biases. Enterprises need to clarify which nodes can be advanced by AI, and which nodes must set human confirmation and decision boundaries.

From this perspective, the focus of AI cost governance is to form a workflow-oriented task layering capability. Strong models, small models, rule-based automation, caching mechanisms and human judgment should not be regarded as single choices that replace one another, but should be recombined within the same workflow. Only when enterprises break processes apart and match different execution methods to different nodes can they avoid two common problems: on one hand, using high-cost models to process a large number of low-value repetitive tasks; on the other hand, suppressing genuinely high-value scenarios worth AI intervention in order to control the bill. Mature AI workflow design needs to realign cost, quality, speed and risk at each task node.

§ iv

Task Boundary Judgment Is Becoming a New Threshold for Enterprise AI Deployment

Although the previous section made an initial classification of enterprise AI tasks, in real deployment, task boundaries are rarely naturally clear. A process often includes multiple links such as context reading, information retrieval, content generation, rule judgment, system calls and human confirmation. For example, even when the task is also “document summarization,” ordinary meeting minutes, customer communication records, contract clause analysis and legal risk warnings have completely different requirements for model capability, cost tolerance and human review. Therefore, task classification can only provide an analytical direction; real deployment decisions still need to return to specific business scenarios, judging node by node whether each task should be handled by a strong model, a small model, rule-based automation or human judgment.

This is also where many enterprises most easily feel confused in AI implementation. Enterprises need to judge three things at the same time: first, whether this task has sufficiently clear business value, and whether AI is saving ordinary operational time or high-value time from key positions; second, what technical path is suitable for this task, whether it is a strong model, small model, RAG, API interface, rule system, caching mechanism or human confirmation; third, after this process runs for a long time, whether the costs of tokens, context reading, tool calls and human review can still be covered by business outcomes. For small and medium-sized enterprises, this problem is especially prominent, because they often do not have enough people who both understand business processes and are familiar with model capabilities, system interfaces and cost structures, making it difficult to independently complete this comprehensive judgment internally.

This is where AI consulting can intervene. What NextAI+ Praxis can help enterprises solve is how to judge which tasks AI should enter, at what cost it should enter, and in which links human confirmation should be retained. Many questions raised by enterprises appear to be “can AI be used for customer service, sales support, contract review or procurement inquiry,” but the deeper question is: in these processes, which consumption can bring real business returns, and which consumption is only handing tasks that could originally be completed by rules, templates or humans to more expensive models.

Therefore, we can say that after enterprise AI deployment enters the agentic stage, maturity depends not only on how many AI tools an enterprise uses, but also on whether the enterprise can judge which task nodes AI should be placed in. Only when the boundaries among models, automation and human labor are concretely clarified can AI consumption be transformed from continuously expanding technical expenditure into a business investment that can be explained, optimized and continuously operated. What NextAI+ Praxis supplements is precisely this judgment capability between business needs, task boundaries, technical paths and cost verification.

§ v

Further Reading

The public research and cases cited in this analysis or considered worth sharing are listed below for further reading (publication years and links are subject to official sources):

  • Microsoft, Announcing the New Work IQ APIs, 2026.
  • Salesforce, Salesforce Introduces New Flexible Agentforce Pricing to Accelerate the Digital Labor Revolution, 2025.
  • Salesforce, Agentforce Pricing, 2026.
  • Longju Bai et al., How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks, arXiv, 2026.
  • Databricks, Introducing AI Spend Controls with Unity AI Gateway, 2026.
  • Axios, Exclusive: Databricks Rolls Out AI Spend Controls, 2026.
  • OpenAI, New Usage Analytics and Updated Spend Controls for ChatGPT Enterprise, 2026.
  • Reuters, OpenAI Introduces Enhanced Usage Analytics, AI Spending Controls for ChatGPT Enterprise, 2026.
  • Financial Times, “We Created a Monster”: Companies Rein in AI Usage as Costs Strain Budgets, 2026.
  • Reuters Breakingviews, Corporate AI Sticker Shock Will Force Restraint, 2026.
Back to Enterprise AI Deployment Signals

Cite as · Enterprise AI Deployment Signals · 28 July 2026

§ Recent signalsBack to Deployment Signals
18 Aug 2026AI being useful is not enough: enterprise AI must first pass the secure deployment gate.04 Aug 2026The new AI stage for cross-border B2B platforms: from connecting transactions to completing transaction capabilities.14 Jul 2026When agents handle enterprise tasks, external interfaces are still stuck in the human internet.30 Jun 2026When AI deployment becomes a product: the logic and applicability boundaries of agent workspaces.16 Jun 2026The next question after AI deployment: can organizational capabilities keep up with AI transformation?

One quarterly digest, no weekly drip.

If you want both columns delivered together, four times a year, in one quiet email — leave an address. Otherwise just bookmark this page.