Value per Person
Intelligence got cheaper and the bill got bigger. That part is fine. The problem is the other side of the ledger.
Token prices fell 98% in three years. Enterprise AI bills went up anyway.
GPT-4-level intelligence that cost about $20 per million tokens in late 2022 now costs about $0.40. Over the same stretch, the average enterprise AI budget went from $1.2M to $7M a year. A simple AI workflow that cost about 4 cents per interaction in 2023 now runs as an orchestrated agent system at roughly $1.20, about 30 times more. One analysis of engineering teams found per-developer token use up more than 18x in nine months.
So the unit price of intelligence collapsed and the bill exploded. That alone is fine. Compute spend growing is what adoption looks like. The problem is the other side of the ledger.
The value side is nearly blank
The numbers from this spring are uncomfortable in sequence. 79% of organizations spending on AI blew through their AI budget in the last 12 months, per a February 2026 survey of 500 finance leaders at companies with more than 1,000 employees. Only 15% of those leaders say they can calculate AI ROI without significant bottlenecks. In the FinOps Foundation's 2026 report, the share of organizations actively managing AI spend went from 31% to 98% in two years, and the single most requested capability is token-level cost monitoring.
Meanwhile only 39% of organizations attribute any EBIT impact to AI at all, and most of those say it is under 5%.
And the most rigorous study we have should be required reading for anyone writing an AI budget. Researchers linked Danish administrative labor records to large adoption surveys and found precise null effects of AI chatbot adoption on earnings and hours, ruling out effects larger than 2%, two years after ChatGPT launched, in a market where most exposed employers had adopted the tools and workers self-reported real productivity benefits. The workers felt faster. The economy could not measure it.
Cost control does not fix this
The instinctive response is discipline: dashboards, budgets, FinOps. Necessary, and not enough. The same survey found that organizations rating themselves very mature on cost management had a higher AI overrun rate, 89%, and the largest mean overspend, 30.9%, versus 69% and 16.1% for beginners. The report's own conclusion: maturity surfaces problems, it does not prevent them.
Token monitoring tells you what AI costs. It tells you nothing about what AI is worth.
The leak is somewhere a cost dashboard cannot see. A person saves 40 minutes with AI, and the 40 minutes dissolves into the workday. Multiply by a thousand employees and you get the Danish result: real micro savings, zero macro impact, and a seven-figure invoice.
The metric that closes the gap
The fix starts with changing the unit of account. AI is bought per token and per seat. Value happens per person, per workflow. So measure value per person, built from two ratios. First, cost per completed task. Tokens spent are irrelevant if the fully loaded cost of the outcome, a resolved ticket, a reviewed contract, a shipped report, is falling. Spending more tokens is often correct. Second, redeployed time. Hours saved are worth zero until they show up somewhere: more volume handled, faster cycle time, people doing higher-order work. If nobody can say where the saved time went, it went nowhere.
Adoption percentage, seats activated, messages sent, tokens consumed: those are activity metrics. They measure motion. The Danish study is what motion without redesign looks like at national scale.
What the high performers actually do differently
One. Put most of the money into people and process. BCG's rule of thumb for AI value is 10-20-70: roughly 10% of the value comes from algorithms, 20% from data and infrastructure, 70% from people, process, and workflow change. Most organizations invert this and spend on technology first, training last.
Two. Redesign the workflow, do not sprinkle the tool. McKinsey's research finds the high performers, the ones reporting real EBIT impact, are about three times as likely to have fundamentally redesigned workflows, the strongest value driver they tested. This is also exactly why the Danish nulls happen. A chatbot bolted onto an unchanged process produces saved minutes that evaporate.
Three. Deploy agents on workflows, not just assistants on people. A chat seat creates value only when a person remembers to use it well. An agent embedded in a process runs every time the process runs, and its output lands in a system instead of a chat window. Active agents in Microsoft 365 grew 15x year over year, and a majority of organizations now run agents on multi-stage workflows rather than single tasks.
Four. Engineer cost per task like you engineer anything else. Route high-volume routine steps to small cheap models and reserve frontier models for judgment-heavy steps. Cache the repeated context. Batch the non-urgent work. Teams that do this cut cost per task several-fold with no quality loss. Cost per task is a design decision, and most organizations have not made it once.
Five. Treat context as infrastructure. The top barriers to agent value are integration with existing systems, 46%, and data access and quality, 42%, far ahead of unclear ROI at 20%. The model is rented and identical for everyone. What is proprietary is the context it operates in: clean processes, documented work, connected data. Every dollar that makes organizational knowledge machine-readable raises the return on every future token.
Six. Manage AI work like management, because that is what it is. The 2026 Microsoft Work Trend Index finds organizational factors drive roughly twice the AI impact of individual factors. The top-performing AI users are distinguished by ordinary-sounding practices: managers set quality standards for AI output, reward people for redesigning work, and make experimentation legitimate. The constraint is not model capability. It is management practice.
Keep yourself honest
The contrarian evidence deserves a permanent seat at the table. An MIT-affiliated analysis last year found 95% of enterprise generative AI pilots produced no measurable P&L impact. A controlled trial found experienced developers were 19% slower with AI assistance on familiar codebases while believing they were faster. Value per person can be negative: low-quality AI output pushed downstream creates work for the receiver. That is the whole argument for measuring outcomes instead of trusting the feeling of speed.
The question for a 2026 budget
Spend will keep rising. Falling token prices keep getting eaten by rising token volume, and that is fine. The risk is walking into 2027 with a seven-figure bill and a shrug where the value line should be.
So skip the question of how much to spend on AI and ask the operational one: for each major workflow, what is the cost per completed task, and where did the saved hours go. Answer that for five workflows and you are ahead of roughly 85% of the market, and you will know exactly which tokens to buy more of.
The organizations that win this decade will spend less per token and get more per person. The gap between those two curves is where the entire prize sits.