
How much do AI agents cost to run?
Key Facts
- Anthropic's Claude Haiku 5.5 costs $0.10/million input tokens under 100k, while Opus 5.5 charges $4/million — a 40x spread according to VentureBeat.
- OpenAI's $200/month plan now includes half the tokens it previously did, effectively doubling cost per token per Yahoo Finance.
- Only 19% of services buyers use outcome-based pricing, despite its promise per CIO Dive.
- Harvey's gross margins fell from 50% to -50% as token usage grew 20x, prompting a switch to open models according to AI Weekly.
- Haiku 5.5 costs 5x more per token above 100k tokens, making request optimization critical per VentureBeat.
- 40% of AT&T's AI workloads now use open models to cut costs per AI Weekly.
Understanding AI Agent Cost Complexity
If you've tried to budget for an AI agent, you've probably discovered the hard part isn't the sticker price — it's that the sticker keeps moving. Agent costs arrive in layers, and each layer prices differently.
The first layer is token-based billing, where you pay per unit of text the model reads and writes. Prices vary enormously by model tier: Anthropic's Claude Haiku 5.5 charges $0.10 per million input tokens for requests under 100,000 tokens, while Claude Opus 5.5 runs $4 per million input and $20 per million output, according to VentureBeat's launch coverage and Anthropic's pricing page. That's a 40x spread for input tokens alone.
The second layer is subscription creep. OpenAI's $200/month plan now includes half the tokens it previously did — effectively doubling cost per token — while a new $500/month tier has appeared, as Yahoo Finance reports. Anthropic's tiers bundle API credits too: $100/month for Max 5x, $200 for Max 20x, and up to $500 shared for Teams.
The third layer is outcome-based pricing, where vendors charge per completed task instead of per seat or token. Zendesk, for instance, charges per resolution only when AI handles an issue 100% end-to-end. But adoption remains rare: CIO Dive reports only 19% of services buyers use outcome-based arrangements, and Gartner calls the trend "more buzz than reality."
Why does this make agent costs so unpredictable? A few structural reasons:
- Request size can change your rate 5x. Haiku 5.5 costs five times more per token above the 100,000-token threshold.
- Usage scales non-linearly. Legal AI firm Harvey watched gross margins fall from 50% to -50% as token usage grew 20x, per AI Weekly.
- Discounts can vanish at scale — Anthropic ends its 15% enterprise discounts once customers exceed contracted volume.
- Vendors expect spend to rise as automation expands, not shrink; Zendesk itself anticipates customers paying more over time.
For a small business, this complexity is the real cost problem. A flat monthly fee for a managed agent — answering calls, following up on leads, handling busywork — sidesteps the token math entirely, which is how we structure billing at Agents by AIQ: predictable, month-to-month, with no usage surprises buried in the invoice.
The practical takeaway from the research is consistent: measure cost per completed task, not just token spend. As Silicon Data CEO Carmen Li puts it, transparent pricing you can forecast "makes budgeting easy" — and that transparency is worth more than the cheapest per-token rate on the market.
Breaking Down Cost Models: Token, Subscription, and Outcome-Based Pricing
The price of running an AI agent depends heavily on which pricing model you're in — and the gap between models can be a factor of 40 or more. Understanding these three structures is the first step to predicting your actual monthly bill.
Per-token pricing: cheap models are getting cheaper, premium models stay premium
Anthropic's Claude Haiku 5.5 charges just $0.10 per million input tokens for requests under 100,000 tokens — a 90% reduction versus Haiku 4.5 — while Claude Opus 5.5 costs $4 per million input and $20 per million output, according to Anthropic's official pricing page. That's a 40x spread on input alone, before you consider that Opus's fast mode runs $8/$40 per million.
Request size matters too. VentureBeat's launch coverage notes the same Haiku model costs 5x more per token above the 100,000-token threshold ($0.50 input / $2.50 output per million). The lesson for anyone operating agents: match the model tier to the task and keep requests lean.
Subscription tiers: shrinking allocations, rising ceilings
Subscriptions look predictable on the surface, but the fine print is shifting. Yahoo Finance reports that OpenAI's $200/month plan now includes half the tokens it previously did — effectively doubling cost per token — alongside a new $500/month tier. Enterprise discounts can evaporate at scale: reporting on AI cost management notes Anthropic's 15% enterprise discounts end once contracted volume is exceeded.
Outcome-based pricing: promising, but rare
Vendors like Zendesk and Pegasystems now charge per resolution or completed case, absorbing underlying AI costs. It's an appealing model for businesses — you pay for finished work, not raw tokens. But adoption is thin:
- Only 19% of services buyers use outcome-based arrangements, per CIO Dive's reporting
- Gartner projects fewer than 25% of tech CEO services contracts will use it through 2031
- Zendesk itself expects customer spending to rise as automation use expands
Gartner analyst Tom Coshow calls the trend "more buzz than reality," advising buyers to ask whether the vendor is genuinely absorbing risk. Seton Hall CIO Paul Fisher adds that outcome pricing complicates contracts because outcomes must be specified precisely.
The practical takeaway
The most useful metric isn't token spend — it's cost per completed task, whether that's an answered call, a resolved ticket, or a booked appointment. Teams that know this number can negotiate rates or shift work to cheaper models when quality allows. Legal AI company Harvey saw gross margins fall from roughly 50% to negative 50% as token usage grew 20x before switching to open-weight models restored profitability.
That's why Agents by AIQ structures agent operations around measurable, completed work rather than opaque token counts — predictable billing is part of the service. Whatever model you choose, scrutinize what a finished task actually costs before signing anything.
Strategic Cost Management for AI Agents
When Harvey, a legal AI company, watched its gross margins collapse from 50% to negative 50% as token usage grew 20-fold, the problem wasn't demand — it was cost control. The fix, a switch to open-weight models, restored profitability and offers a playbook any business running AI agents should study.
The first step is measuring the right number. Token spend tells you what you consumed; it doesn't tell you what you accomplished. Teams that track cost per completed task — an answered call, a resolved ticket, a booked appointment — can negotiate better rates or move work to a cheaper model when quality allows, according to reporting on AI cost management. That number is also the basis for any outcome-based pricing conversation with a vendor.
Model selection is the second lever, and the pricing gaps are enormous. Anthropic's Claude Haiku 5.5 runs $0.10 per million input tokens, while Claude Opus 5.5 charges $4 per million input — a 40x difference for the same unit of work. The practical approach, which Pegasystems applies by selecting cost-effective models per task, is routing repetitive work like retrieval and classification to small models and reserving premium models for complex reasoning. Rogo's Alex Alexander Wang notes Haiku-class models are now "fast and cheap enough that we can run it a lot."
Request optimization is the third, often-overlooked lever. Haiku 5.5 costs 5x more per token above the 100,000-token request threshold, so structuring agent requests to stay under pricing thresholds directly cuts effective costs. Cache usage matters too: Haiku cache reads cost $0.01 per million versus $0.125 for writes, making repeated context worth caching.
A few practical controls to put in place:
- Set spend alerts and hard limits on your API accounts — OpenAI lets you cap usage and return errors when limits are hit, separate from its own tier limits.
- Budget for full-price usage at scale, since enterprise discounts (Anthropic's run 15%) can end once you exceed contracted volume.
- Review subscription tiers annually — OpenAI's $200/month plan now includes half its previous token allocation, effectively doubling per-token cost.
- Scrutinize per-resolution vendor pricing carefully; Gartner analyst Tom Coshow calls the outcome-based trend "more buzz than reality" and advises confirming the vendor actually absorbs risk.
This is the discipline we apply when operating agents for small and mid-size businesses at Agents by AIQ: match each task to the right model tier, monitor usage fees against subscription tiers, and keep the billing model predictable month to month. Carmen Li, CEO of Silicon Data, puts the goal plainly — transparently priced services make budgets easy to forecast. With token prices falling for small models but subscription allocations shrinking, that forecastability has to be engineered, not assumed.
Frequently Asked Questions
How much does it actually cost to run an AI agent per month?
Why are AI agent costs so hard to predict?
Is outcome-based pricing (paying per completed task) a good deal?
Are AI subscription plans getting cheaper as token prices drop?
How can I keep my AI agent costs under control?
Is there a way to avoid the token math entirely?
The Hidden Costs of AI Agents: What Businesses Need to Know
Running AI agents involves navigating a complex web of pricing models, from fluctuating token costs to subscription tiers and rare outcome-based arrangements. The key takeaway? Measuring 40x price disparities between model tiers and optimizing request sizes can drastically impact expenses. For small and mid-size businesses, unpredictable costs often outweigh the benefits of raw AI capabilities. By focusing on cost-per-completed-task—whether it’s a resolved ticket or a booked appointment—teams gain control, enabling smarter model choices and budget predictability. At Agents by AIQ, we design billing around measurable outcomes, eliminating the guesswork of token math. If you’re managing AI agents, start by auditing your current spend, matching tasks to cost-effective models, and setting spend controls. The goal isn’t just to cut costs—it’s to align AI investment with tangible business results. Ready to simplify your AI operations? Book a call to explore how we can tailor an agent solution that fits your needs without the hidden fees.