Null StudioNullStudio

Blog · October 2, 2026 · 9 min read

AI Cost Control: How to Keep LLM Costs Predictable in Production

By the Null Studio team

TL;DR: AI features are the first software most companies have shipped where the running cost moves with every click. You keep that cost predictable the same way you keep any variable cost predictable: measure it per unit of business value, not per month; design so that expensive model calls happen once instead of on every page view; send the model only the context the task needs; route the easy majority of requests to a small model; and put hard limits in the code, per user, per feature and per agent run, so that a bug or a single heavy user cannot turn into a surprise invoice. None of this needs a cheaper vendor. It needs the cost designed in from the first sprint.

Most AI cost problems are not discovered in the spec. They are discovered on the first full invoice after launch, when a feature that cost almost nothing in the demo turns out to be called far more often than anyone assumed, with far more context than anyone intended.

The response at that point is usually to look for a cheaper model or a better rate. Sometimes that helps. More often the bill is high because of how the feature is built, and a cheaper model only lowers the price of a design that is still wasteful. This article is about the design.

We have touched on this from other angles in choosing an AI model and adding AI to an existing product. Here we go deeper on the one question both of those leave open: once it is live, how do you stop the cost from running away?

Why AI cost behaves differently from the rest of your stack

Most software costs are close to fixed. A server, a database and a set of SaaS licenses cost roughly the same whether your users are busy or quiet. Hosted AI models charge per request, priced by how much text (or audio, or image) goes in and how much comes out. That has three consequences that catch teams out.

Usage is uneven. Your heaviest users do not cost slightly more than your lightest ones. They can cost many times more, because they trigger the feature more often and usually on bigger inputs. An average cost per user hides exactly the accounts that hurt.

Input size is invisible. Nobody sees the prompt. A feature that sends a whole document, the whole conversation history and a long set of instructions on every request looks identical in the interface to one that sends a paragraph. The difference only shows up on the bill.

Loops multiply. An agent that calls a model, reads the result, then calls it again can make many calls for one user action. A retry bug or an agent that never decides it is finished can make thousands.

None of these are reasons to avoid AI features. They are reasons to treat model calls as a metered resource from day one, the same way you would treat SMS or payment fees.

Measure cost per unit of value, not per month

A monthly AI bill tells you almost nothing. The number that matters is cost per unit of the thing your business actually sells or saves: per call handled, per document processed, per lecture summarized, per active user per month.

That means logging every model call with enough attached to answer "what was this for": which feature, which customer, which model, how many tokens in and out, and whether the result was used. This is a small amount of engineering, and without it every later decision is a guess.

With it, you can answer the questions that matter:

On voice work such as the agents we built for CallGuard AI and CallSetter AI, the natural unit is the call minute, because that is how the business thinks about its own economics. On our own products, LectureNotes AI and Lifemaxxing AI, the natural unit is the active user, because that is what a subscription pays for. Pick the unit that matches revenue and you can see at a glance whether a feature pays for itself.

Design the expensive call to happen once

The single biggest saving in most AI products comes from moving model calls off the read path.

Generate on change, not on view

If a record gets a summary, generate that summary when the record changes and store it. Do not generate it every time someone opens the page. A summary that is read many times and written once should cost one call, not one per view. The same applies to classifications, tags, extracted fields and embeddings for search.

Cache what repeats

Many requests are identical or nearly identical: the same question about the same help article, the same instructions at the top of every prompt. Exact-match caching of results is simple and safe. Most major providers also offer some form of prompt caching, which makes a repeated long prefix (your instructions, a product manual, a policy document) cheaper on subsequent calls. Structure prompts so the stable part comes first and the variable part comes last, and you get the benefit with no change in behavior.

Batch what is not urgent

Work that does not need an answer in the next few seconds, such as overnight reprocessing, bulk classification or backfilling summaries for old records, can often go through a provider's batch interface, which is usually priced below the real-time one. It is a scheduling decision, not a model decision.

Send less

Context is the quiet cost driver. Teams pass in everything because it is easier than deciding what the task needs, and because the model copes. It copes expensively.

Three habits keep context in check:

  1. Retrieve, do not dump. Pull the few relevant passages from a document or knowledge base rather than the whole thing. This is the core of a well-built internal AI knowledge base, and it usually improves answers as well as cost.
  2. Trim conversation history. Long chats do not need every earlier turn resent in full. Summarize older turns or keep only what is still relevant.
  3. Cap the output. Ask for the length you need. A model asked for "a summary" may write far more than the interface can usefully show.

Smaller inputs are also faster, which matters a great deal in anything real-time, such as a voice agent where the caller is waiting.

Route the easy majority to a small model

In most products the requests are not equally hard. A large share are routine: classify this message, extract these fields, answer a question that has a clear answer in the provided text. A smaller, cheaper model often handles those well. A smaller share genuinely need a frontier model.

The pattern we use is to decide on the cheapest model that passes your own evaluation set for each task, and send each task to that model, with an escalation path when the small model signals it cannot do the job. That only works if you can switch models without rewriting the feature, which is why we build a thin layer between the product and the provider, as described in choosing an AI model.

Sometimes the right answer is no model at all. If a rule, a lookup or a regular expression gets the job done reliably, it is cheaper, faster and easier to test than any model.

Put hard limits in the code

Monitoring tells you about a problem after it happens. Limits stop it from becoming expensive while it happens. Every production AI feature should have them.

Per user and per plan

Set a usage allowance for each plan and enforce it in the product. That protects margins from the heaviest accounts and gives sales a clean upgrade path. How those limits should be packaged is a pricing decision, and we cover it in SaaS pricing and packaging. The point here is that the limit has to exist in code, not only on the pricing page.

Per agent run

Any agent that can call a model repeatedly needs a ceiling on steps, tokens and wall-clock time for a single task, after which it stops and reports where it got to. This is the cost version of the containment rules in AI agents with write access: an agent that cannot do damage can still run up a bill.

Per feature and per account

Set budgets with alerts at the provider level and inside your own logging, so that a spike in one feature or one tenant is noticed in hours, not at month end. Decide in advance what happens when a budget is hit: degrade to a smaller model, queue the work, or switch the feature off with a clear message. A decided fallback is far better than an outage decided under pressure.

Against abuse

Public-facing AI features attract people who want free access to a model. Rate limits, sign-in requirements for expensive actions and basic checks on input size stop a single bad actor from becoming your largest customer.

When the numbers still do not work

Sometimes a feature is well built and still costs too much relative to what it earns. At that point the options are commercial rather than technical: charge for the feature separately, include it only on higher plans, set a usage allowance with paid top-ups, or narrow the feature to the cases where it creates the most value. It is better to make that call before launch, from a cost model built on realistic usage, than after customers have come to expect the feature for free.

The opposite also happens. Teams sometimes cut a valuable feature because the bill looks large in isolation, when the cost per unit is small against what the unit earns. Cost per unit of value settles that argument as well.

Questions to ask before you commission an AI feature

Whether you build in-house or bring in a studio, these questions show quickly whether cost has been designed in or left for later:

  1. What is the expected cost per unit of value (per user, per call, per document) at realistic usage a year from now?
  2. Which model calls happen on read, and could they happen once on write instead?
  3. How much context goes into a typical request, and why does each piece need to be there?
  4. Which tasks could a smaller model or plain code handle, and how was that tested?
  5. What hard limits exist per user, per plan and per agent run, and what happens when one is hit?
  6. Can we see cost by feature and by customer, not only as a monthly total?
  7. How hard would it be to switch model providers if prices or quality change?

A vendor who answers these with specifics has thought about running the product, not only demoing it.

The bottom line

AI cost is controllable, but only if it is treated as a design constraint from the first sprint rather than an invoice to be negotiated later. Measure cost against the unit your business earns on, move expensive calls off the read path, send less context, route routine work to small models and put real limits in the code. Done that way, a growing AI bill means a growing product, not a growing problem. That discipline is part of how we keep builds fast without leaving clients with surprises, as set out in the AI-native studio playbook.


Worried about what an AI feature will cost at real volume, or already looking at a bill that grew faster than usage? Book a demo and we will walk through your cost per unit, where the spend is going and which changes would move it most. See our AI work: voice and intake systems for CallGuard AI, CallSetter AI and Fortell, plus our own products LectureNotes AI and Lifemaxxing AI.

FAQ

How do I reduce the cost of an AI feature without making it worse?

Start with design rather than vendor rates. Generate results once when data changes instead of on every page view, cache repeated requests and long instruction prefixes, retrieve only the relevant context instead of whole documents, and route routine tasks to a smaller model that passes your own evaluation set. These usually cut cost without any loss in quality, and often make the feature faster.

How do I stop an AI agent from running up a huge bill?

Give every agent run a hard ceiling on steps, tokens and time, after which it stops and reports where it got to. Add per-user and per-account limits, budget alerts at the provider and in your own logs, and a decided fallback for when a budget is hit, such as switching to a smaller model, queuing the work or pausing the feature.

How should I measure AI costs?

Per unit of business value rather than per month: cost per call handled, per document processed or per active user. Log every model call with the feature, customer, model and token counts attached, so you can see which feature and which accounts drive the spend and compare it with what that unit earns.

Want it built, not just explained?

We design, build and run these systems end-to-end — shipped in days, not months.

Book a demo →

Keep reading