AI Inference Costs: Gartner Fivefold 2028 Analysis
Key Takeaways
- Measure cost per completed agent workflow, not average cost per prompt.
- Add routing, caching, token limits and approval gates before agent usage scales.
- Require business value and ownership for agent projects, or cancellation risk rises.
Agent workflows are moving cost control from the model console to product, finance and governance meetings.
The cheapest AI demo is still the one that never meets a real user. Once an agent starts planning, retrieving context, calling tools, checking its own work and doing it again because the first answer was not usable, the invoice stops looking like a chatbot experiment. Gartner’s new forecast puts a number on the unease: inference cost per agentic workflow is expected to increase more than fivefold through 2028. That is not a compliance deadline, but it should be treated like one by anyone putting autonomous workflows into production.
What Gartner is actually warning about Gartner’s newsroom release says AI
inference costs per agentic workflow will increase more than fivefold through 2028. Techstrong.ai, reporting on the same Gartner research, adds the awkward part: the increase is expected even as token prices fall, because more capable reasoning models can consume more tokens and compute as agents work through multiple steps and decisions. In plain English, cheaper ingredients do not help much if the recipe now uses far more of them. This matters because an agentic workflow is not a single prompt with a tidy response. Techstrong.ai describes agents working through multiple steps and decisions, which is precisely where finance teams should stop accepting average cost per query as the relevant metric. The unit of account is the completed workflow, including failed attempts, retries, retrieval, tool calls and escalation. If your dashboard cannot show that, it is not measuring the thing Gartner is warning about. The practical translation is simple. Product teams need to decide which workflows deserve agentic treatment, engineering teams need to route cheaper work to cheaper models, and finance needs limits before usage scales. Calling this an engineering concern is how surprises become quarterly explanations.
The cost control stack is now governance
Techstrong.ai says Gartner recommends inference tiering, token limits and monitoring to keep agentic AI costs under control. Zylos Research describes a related discipline, Inference FinOps, focused on governing, routing, caching and arbitraging AI compute spend across providers. Strip away the label and the obligation is familiar: know what is running, why it is running, who approved it, and when it should stop. For builders, that turns architecture choices into policy controls. Routing decides whether a request needs a frontier reasoning model or a cheaper model. Caching decides whether the system pays again for context it has already fetched. Approval gates decide whether the agent can keep spending tokens, invoke tools or hand off to a human. These are not decorations for an architecture review. They are the cost equivalent of access controls. Zylos Research also says inference accounted for 85% of the enterprise AI budget in 2026 and roughly two thirds of global AI compute spend. It estimates a single chatbot API call might cost $0.001, while a multi step agent that plans, retrieves context, invokes tools, reflects and self corrects can cost $0.10 to $1.00 per task completion. Those figures will not map cleanly to every vendor contract, naturally. They do explain why procurement will start asking harder questions than whether the demo looked fluent.
The cancellation risk is not about model magic Forbes, citing
Gartner, reports that over 40% of agentic AI projects may be canceled by the end of 2027, attributing cancellations to management issues rather than model capabilities. The cited reasons are poor governance, undefined business value and insufficient operational discipline, including cases where chatbots are mislabeled as true agents. That is the less glamorous version of the story, which usually means it is the one worth reading. This is where law and policy habits help, even when no regulator is forcing the issue. Define the purpose before deployment. Record the decision rights before the system acts. Set thresholds for cost, error handling and human review before the pilot becomes part of daily operations. If that sounds like paperwork, congratulations, you have discovered production. Inside Privacy reports that the UK Information Commissioner’s Office published early thinking in January 2026 on the data protection implications of agentic AI, while noting that the report was not formal regulatory guidance. The distinction matters. The ICO did not create a new rulebook in that report, but its interest signals that autonomy, data access and accountability are now policy questions as well as product questions.
What changes in practice through 2028 Gartner’s
2028 forecast does not require anyone to shut down agent projects. It does require a different approval conversation. A workflow should have an expected value, an expected completion cost, a fallback path and a named owner before it gets broad access to users or internal tools. If those four items are missing, the agent is not autonomous. It is unsupervised spending with a friendly interface. The next useful move is not to ban agents or buy another dashboard because a vendor used the word governance. Start with the workflows where autonomy has a measurable payoff, then apply routing, caching, token limits, monitoring and approval gates before volume arrives. Watch Gartner’s cost forecast through 2028, the reported cancellation rate through the end of 2027, and the privacy regulators circling around agentic systems. Builders who treat inference economics as product policy will have fewer dramatic meetings later. Their lawyers may even stop using cheerful phrases in press statements, which is how you know things are improving.
