We observed three questions that have run the AI conversation through the first half of 2026. Is adoption generating ROI? Should companies push usage or cap it? Will open-source models replace the frontier?
Each one starts too far downstream. We believe the more useful management question is: What combination of people, software, and models can produce the required result at the right level of quality and the lowest sustainable cost?
For the 2026 State of AI report, we surveyed roughly 300 executives at software companies building AI products, interviewed finance leaders who own AI budgets, and studied what companies actually pay for. Our conclusion is that AI economics is an operating-design problem before it is a spend-management problem. Tokens, licenses, and benchmark scores are inputs; the unit of value is the workflow outcome.
Companies should identify the workflows AI can materially redesign, set the quality and risk bar, and only then optimize the mix of labor, software, and models. A token cap or model preference should never come first.
AI activity is high. Deep use is rare.
The gap between AI pilots and measurable P&L impact is generally familiar by now. Many pilots begin with a horizontal tool and a broad mandate to "use AI." The result is often activity without any change in how work moves through the company.
Our survey also shows the same gap. At the companies we surveyed, 71% of R&D employees have access to AI tools, 62% use them daily, and only 39% are power users. The drop-off is steeper in sales, marketing, and corporate functions. Most companies have bought AI access. Fewer have established deep, repeatable use.

More usage is not automatically more valuable. Once AI is widely available, employees default to it even for work where it adds a step rather than removing one. An AI-generated first draft still gets rewritten. A human repeats the analysis to verify the output. Two teams may unknowingly use AI to solve the same problem. And when relevant context is scattered across CRM, ERP, ticketing, knowledge-management, and communication systems, employees can spend more time finding, copying, and reconciling information than the AI ultimately saves.
These costs often stay hidden because they rarely appear in the AI budget. They show up as rework, duplicate effort, extra review, slower handoffs, and inconsistent decisions. A task finished faster in one application can still make the end-to-end process slower.
This creates two distinct failure modes: too little experimentation to find valuable use cases, and indiscriminate use that layers AI onto existing work without removing anything. Companies need to give employees enough freedom, and enough spend, to test new approaches and find where AI changes the work. But they also need to watch the full workflow and separate AI that eliminates effort and AI that merely shifts or duplicates it.
That is why the cost often appears before the ROI does. A license or API bill can be counted immediately. Time saved, faster cycle times, better conversion, or lower support costs may only become visible when a company changes the workflow, sets a baseline, and identifies the business metric that should shift.
In our view, some of the best candidates for redesign share four traits. The work is high-volume and repeatable, with digital inputs and outputs. It has a clear quality threshold and an observable business result. It is constrained by queues, handoffs, or routine review rather than by irreplaceable judgment. And it carries enough economic weight to justify the integration, monitoring, and change management that redesign requires.
For those types of workflows, AI can become the default path and human staffing the exception. For ambiguous, low-frequency work, a copilot may still improve individual productivity, but it should not carry the same ROI expectations (at least for now).
Ramp's Deal Desk is a good example. A 2-person team runs it, managing thousands of contracts per month for 600+ sellers at a business with over 70,000 customers. The team has redesigned the core workflows so AI handles the standard path: intake, pricing, other terms, document generation, and signature routing. This allows the operators to focus on the small subset of deals that require judgment -- complex negotiations, custom terms, and more.
The lesson is to give teams room to find where AI changes the work, then concentrate investment on the workflows where automation can remove the most human touch and handoffs while preserving quality. In practice, that could look like a short, time-boxed workflow review once a use case shows repeated demand: establish a baseline for volume, cycle time, rework, and cost; map the steps and handoffs; define the quality bar and exception rules; then pilot the redesigned process against the baseline. Workflows that show a material improvement earn deeper integration and change-management resources; those that do not remain lightweight copilots or are retired.
Metering token spend is managing the wrong number
According to Ramp's business spend data, the top 1% of companies spend about $7,500 per employee per month on AI. The top 10% spend about $630. The median company spends just $12.

That spread does not tell us which company is better managed. It tells us that companies are operating different workflows, at different levels of maturity, and with vastly different outcome economics. In our view, an average per-employee target would be mostly meaningless.
The current debate generally treats token use as either evidence of ambition or as a cost to suppress. Both are weak proxies. Token counts are easy to meter, but they do not tell a manager whether the company approved more contracts, resolved more cases, shipped better software, or closed its books faster.
Metering still matters. Models are variable-cost infrastructure, multi-step agentic workflows can burn 5 to 30 times as many tokens as a simple chat exchange, and 68% of the builders we surveyed said the true cost of internal AI is very or moderately difficult to predict. But metering is necessary plumbing. A cap can prevent a runaway bill; it cannot determine which workflows deserve investment.
Our survey also shows how sharply usage scales with company size. Only 30% of companies with 50 or fewer employees spend $50,000 or more a month on internal AI tokens. Among companies with 1,000 or more employees, 36% spend $1 million or more, and 14% that spend more than $5 million.

Generally, more disciplined companies separate activity from value. Token counts are easy to measure, but they become a poor target the moment teams are rewarded for maximizing them. Stronger systems charge AI usage back to cost centers and connect it to outcomes such as features shipped, cases resolved, or cycle time reduced. One early-stage company gives engineers unlimited tokens, but holds managers accountable for three to five times the productivity in return.
Model choice follows the workflow
Once the workflow and quality bar are defined, the open-versus-frontier debate becomes a routing decision rather than a company-wide ideology.
Frontier models dominate enterprise wallet share. In our latest survey, 82% of companies surveyed use Anthropic and 71% use OpenAI, compared with roughly 4% to 7% using open-source models from DeepSeek and Moonshot AI.

At the same time, the average number of model providers per company rose from about 3.1 to 3.3 in the last six months and most builders use a combination of closed and open-source models. New releases, including Moonshot AI’s Kimi, alongside models from DeepSeek, Qwen, Mistral, and others, are expanding the range of tasks that can be served by lower-cost or customizable alternatives.
Running more providers is deliberate portfolio architecture. The rule is simple: use the lowest-cost model that reliably clears the quality bar for a task, and escalate only when the added capability is worth more than the added cost.
The resulting architecture is a barbell. Open, customized, or lower-cost models handle routine, high-volume work; frontier models handle complex reasoning, exceptions, and high-stakes decisions. Open source models tend to win volume. Frontier models often win premium work.
One operating envelope for people, software, and models
Today, many companies budget headcount, software, and model usage separately. However, work does not happen separately.
Optimizing those line items independently can make the total system worse: cap model usage and add manual reviewers; cut a software tool and create more coordination work; freeze headcount and buy technology without redesigning the process.
AI pushes teams toward a different unit of management: cost per outcome. A deal-desk leader, for example, should own the total cost of moving a contract from submission to compliant approval at an agreed turnaround time. Within that envelope, the team can rebalance among reviewers, workflow software, and model usage.
That envelope must include the costs that line-item budgets hide: rework, escalations, latency, human oversight, vendor concentration, security, and change management. The cheapest model is not cheaper if it creates more reviews; software is not efficient if it adds another layer without removing work; and headcount is not necessarily more flexible when volume is volatile.
Governance changes too. Finance sets the envelope and return hurdle; workflow owners own quality, service levels, and cost per outcome; platform teams provide routing, observability, and security. AI becomes part of the operating budget, not a special pool.
Coinbase has described the direction of travel. It has said it cut AI spend in half while token usage continued to grow, using routing, caching, and lower-cost open models as defaults. The lesson is not to minimize tokens. It is to engineer the system so more useful work fits inside the same or a lower operating envelope.
Fund discovery and budget outcomes
We believe the answer collapses the three debates into one operating model. Let teams explore widely enough to find valuable patterns, then redesign and fund the workflows that prove out. Discovery and discipline work together.
Give people room to experiment, compare tools, and surface repeatable patterns. Graduate the strongest patterns into workflows with scale, clear rules, measurable outcomes, and handoffs worth removing. Redesign those end to end, put people on the exceptions and the judgment, and hold the team to one envelope across headcount, software, and models. Route each task to the lowest-cost model that clears the quality and risk bar.
By 2027, we believe AI advantage will not come from spending the most or capping the hardest. It will come from operational clarity: knowing which work to redesign, what the outcome must achieve, and how to allocate people, software, and models to deliver it at the lowest sustainable cost.
Published:
August 13, 2026



