AI cost management: how to understand and control what your business spends on AI
The short answer
AI cost management means knowing what your business spends on AI, where that money goes, and how to get the same work done for less. It covers two kinds of spend: seats (like ChatGPT Business or Claude Team) and usage (the tokens your agents and API calls consume). The price of AI per token falls every year, yet most AI bills go up, because agents read and write far more tokens than chat ever did. The biggest savings come from matching the model to the job, not paying to resend the same context, and seeing which teams and agents actually drive the spend.
What the main AI models cost (per 1 million tokens)
| Model | Input | Output | Cached input | Good for |
|---|---|---|---|---|
| OpenAI GPT-6 Astra | $10.00 | $50.00 | $1.00 | The hardest reasoning and research tasks |
| OpenAI GPT-6.1 Sol | $2.00 | $10.00 | $0.10 | Most everyday business work |
| OpenAI GPT-6 Luna | $0.10 | $0.50 | $0.01 | High-volume, simple tasks (sorting, tagging, short summaries) |
| Anthropic Claude Opus 5.5 | $4.00 | $20.00 | $0.20 | Complex writing, analysis and agent work |
| Anthropic Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 | Most everyday business work |
| Anthropic Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | Fast, lighter tasks |
| Google Gemini 3.1 Pro (preview) | $2.00 | $12.00 | $0.20 | Long documents and multimodal work (price shown is for prompts up to 200K tokens) |
| Google Gemini 3.8 Flash | $0.75* | $3.75* | $0.075* | Fast, low-cost general work |
| Google Gemini 3.1 Flash-Lite | $0.25 | $1.50 | $0.025 | High-volume, simple tasks |
API list prices from each provider, checked October 1, 2026. *Gemini 3.8 Flash is at an introductory price through December 31, 2026, then $1.50 input and $7.50 output. A token is roughly three-quarters of a word. Prices change often, so check the sources below before you budget.
What is AI cost management?
Most businesses now pay for AI in two different ways, and they're usually tracked in two different places.
- Seats. A flat monthly price per person for tools like ChatGPT Business (from $20 per user per month on annual billing) or Claude Team ($20 per seat per month on annual billing, or $100 for a premium seat). Easy to budget, easy to forget about.
- Usage. Pay-as-you-go spend on model APIs, measured in tokens. This is what agents, automations and custom tools run on. It's where costs surprise people, because it scales with how much work the AI does, not with headcount.
AI cost management is the habit of tracking both, tying the spend back to the teams and workflows that create it, and making deliberate choices about which model does which job.
Why AI bills go up while AI prices go down
The price of AI is falling faster than almost any technology before it. Epoch AI found that the cost of reaching a given level of model performance has dropped by between 9x and 900x per year, depending on the task. So why does the invoice keep growing?
Because the way we use AI changed. A chat answers one question. An agent works in a loop: it reads instructions and background documents, calls a tool, reads the result, decides what to do next, and repeats. Every step usually re-reads everything that came before.
Here's a simple example. Say an agent has 40,000 tokens of context (instructions, a contract, an email thread) and takes 10 steps to finish a task. That's about 400,000 input tokens for one run. On a model priced at $2 per million input tokens, that's roughly $0.80. Run it 50 times a day and it's about $40 a day, or close to $900 a month, for one workflow. Most of that money pays for the model to re-read the same context, again and again.
The fix isn't to stop using agents. It's to stop paying for the waste.
How much does AI cost per month?
For a small team, a realistic picture looks like this:
- Seats: roughly $20–25 per person per month for standard business plans, and $100 or more for heavy users on premium seats. A 10-person team on standard seats is about $200–250 a month.
- Usage: anywhere from a few dollars to thousands, depending on how many agents and automations you run and which models they use.
Model choice matters more than most people expect. Summarizing a 30-page document (about 15,000 tokens in and 1,000 out) costs about $0.04 on Claude Sonnet 5.5 or GPT-6.1 Sol, and about $0.002 on GPT-6 Luna. That's a 20x difference for a task the smaller model can often handle just as well.
7 ways to cut AI costs without cutting what AI does
1. Match the model to the job
Use the big models for hard reasoning, and small models for sorting, tagging, extracting and short summaries. Many teams default to the most capable model for everything, which is like sending every package by overnight courier.
2. Stop paying to resend the same context
Most providers discount repeated input heavily. Cached input on GPT-6.1 Sol costs $0.10 per million tokens instead of $2.00, and Claude's cache reads cost a tenth of normal input. In the example above, caching the shared context would take that $0.80 run closer to $0.10. The bigger win is architectural: give agents one organized source of company context they can pull the relevant pieces from, instead of pasting the same documents into every prompt and every tool.
3. Keep outputs short
Output tokens cost about five times as much as input on most models. Ask for the answer, not an essay, and set length limits on automated tasks.
4. Put limits on agent loops
An agent that gets stuck can retry for a long time. Set a maximum number of steps and a budget per run, and look at runs that fail. A failed run costs money and produces nothing.
5. Know who's spending what
You can't cut what you can't see. Break spend down by team, by agent and by workflow. Often a single automation turns out to drive most of the bill, and once you can see it, it's usually easy to fix.
6. Audit seats every quarter
AI tools spread fast. It's common to find people paying for overlapping tools, or seats nobody has opened in a month. Consolidating on fewer tools also makes the next step easier.
7. Recheck prices regularly
Prices move every few months, usually down, sometimes up (Gemini 3.8 Flash's introductory price, for example, doubles in January 2027). A quick monthly check can move a workflow to a cheaper model of the same quality.
The cheapest model isn't always the cheapest
One honest trade-off: a small model that gets the task wrong and retries three times can cost more than a bigger model that gets it right the first time, and wrong answers have costs of their own. The goal is the lowest cost for work you can trust, not the lowest price per token. Test the cheaper model on real examples before switching a workflow over.
Where AgentOS fits
AgentOS is the AI foundation for your business: it keeps your company's context, tools and skills in one place and shares them with every AI agent your team uses, including ChatGPT, Claude and Cursor. That helps with cost in a few practical ways:
- Context once, not everywhere. Agents pull the relevant company context over MCP instead of each tool needing its own copy, or a person pasting the same documents into every chat.
- Skills make work repeatable. A written-down way of doing a job means fewer retries and less trial and error per run.
- Approvals before important actions. Consequential writes wait for a person, so a confused agent doesn't create expensive problems downstream.
AgentOS has a free plan (200 credits a month, no card), Pro at $99 a month and Team at $299 a month for three people. See pricing, or read what AgentOS is. If your team uses several AI tools side by side, AgentOS vs. ChatGPT and AgentOS vs. Claude explain how they work together.
Frequently asked questions
How much does AI cost per month for a small business?
Most small teams pay $20–25 per person per month for business AI seats like ChatGPT Business or Claude Team, plus usage for any agents or automations. A 10-person team on standard seats is about $200–250 a month before usage. Usage can range from a few dollars to thousands, depending on how much work agents do and which models they run on.
What is cost per token?
AI models charge by the token, roughly three-quarters of a word. Prices are quoted per million tokens, separately for input (what you send) and output (what the model writes). Output usually costs about five times more than input, and repeated, cached input is usually much cheaper.
What is the cheapest AI model?
As of October 2026, among the major providers, OpenAI's GPT-6 Luna ($0.10 input, $0.50 output per million tokens) and Google's Gemini 3.1 Flash-Lite ($0.25 input, $1.50 output) are among the lowest-priced. They're great for simple, high-volume tasks, but test them on your real work before relying on them for anything complex.
How do I reduce LLM costs?
Match the model to the task, use prompt caching for repeated context, keep outputs short, cap agent steps and budgets, and track spend by team and workflow. Giving agents one organized source of company context, instead of resending documents in every prompt, is often the biggest single saving.
Is a ChatGPT or Claude subscription cheaper than the API?
For people chatting with AI during their workday, seats are usually simpler and predictable. For agents and automations that run on their own, you pay for usage through the API, and cost depends on volume and model. Many businesses need both.
Why are my AI costs going up if prices are falling?
Because usage grows faster than prices fall. Agents work in loops and re-read their context at every step, so a single task can use hundreds of thousands of tokens. Price per token drops every year, but tokens per task have grown even faster.
Give every AI tool your team uses the same context. Start free with 200 credits a month. No card required.
Set up your workspaceSources
- Epoch AI: LLM inference prices have fallen rapidly but unequally across tasks
- OpenAI API pricing
- Latent Space: OpenAI DevDay 2026 (GPT-6.1 Sol pricing)
- BenchLM: OpenAI API pricing (Sep 30, 2026)
- Claude pricing (API and Team plans)
- Google Gemini API pricing
- OpenAI Help: ChatGPT Business pricing and seats
- AgentOS pricing
AgentOS is built by Devcore. Found something out of date on this page? Tell us at tryagentos.net and we'll fix it.