
A Solutions Architecture Approach to Getting Them Under Control
If you’ve started using AI tools in your business — anything from an LLM API embedded in your product to a handful of AI assistants across your team — there’s a good chance the bill has started climbing faster than anyone expected. It’s becoming one of the most common conversations we’re having with tech-focused SMBs right now: AI spend crept up gradually, nobody quite knows why, and the instinct is to assume you just need to “use it less” — which, much like cloud spend a few years ago, usually isn’t the real answer.
AI cost problems are rarely about usage being inherently too high. They’re almost always about adoption that happened faster than the architecture around it — because early on, when it was one team experimenting with one tool, it didn’t need a plan. The good news is that this is a genuinely fixable problem, provided you’re solving the right one.
Why AI costs spiral faster than cloud costs did
AI spend has a few characteristics that make it spiral even quicker than the cloud cost creep many businesses have already been through:
- Per-call pricing with no natural ceiling.
- Unlike a fixed server bill, AI API costs scale directly with usage — every request, every token, every generation adds up in real time, often with nobody watching the meter.
- Model choice matters enormously, and defaults are usually the most expensive option.
- It’s common to see a task running on a large, expensive model when a smaller, cheaper one would do the job just as well — because nobody revisited the choice once it was working.
- Sprawl across tools and teams.
- AI capability shows up in a chatbot subscription, an embedded feature in your CRM, a coding assistant, an API integration — each billed separately, each decided on its own, with no one holding the full picture.
- Experimentation that never gets reviewed.
- A proof of concept that worked gets left running in production. Nobody circles back to check whether it’s still needed, or whether it was ever optimised for cost.
- Redundant calls and poor caching.
- The same prompt, or something close to it, gets sent to the model repeatedly because there’s no caching or deduplication layer — paying full price for an answer you already had.
None of this is exotic, and none of it is a sign of anyone doing something wrong. It’s the predictable result of a genuinely useful technology being adopted quickly, by teams focused — rightly — on what it can do, not yet on what it costs to run at scale.
Why “just use it less” isn’t the right frame
The instinctive response to a high AI bill is to restrict usage — cut licences, limit API calls, tell the team to be more careful. That can help at the margins, but it treats the symptom, not the cause, and it risks throttling the exact capability that made the tool worth adopting in the first place.
A Solutions Architecture approach starts from a different question: not “how do we use this less?” but “does how we’re using this actually match the value it’s creating?” That reframing tends to surface bigger, more durable savings than simply asking people to be more frugal.
Where the real savings usually are
In our experience, the biggest and most durable wins tend to fall into a small number of categories:
- Right-sizing the model to the task.
- Not every task needs the largest, most capable model available. Matching model choice to actual task complexity — using a smaller, cheaper model where it performs just as well — is often the single biggest lever available, and one of the easiest to overlook.
- Caching and deduplication.
- Where the same or similar requests are being made repeatedly, caching results rather than regenerating them can cut costs substantially, particularly for high-volume, low-variability use cases.
- Prompt and workflow efficiency.
- Poorly structured prompts, unnecessary context, or overly chatty multi-step workflows all add cost without adding value. Tightening these up is usually low-risk and high-impact.
- Consolidating tools and licences.
- AI capability adopted piecemeal across different teams often duplicates itself — several tools solving overlapping problems, billed separately. Consolidating onto fewer, well-chosen tools reduces both cost and the operational overhead of managing them.
- Visibility and ownership.
- Just as with cloud spend, you can’t optimise what you can’t see. Attributing AI spend to the team, product, or feature responsible — rather than watching one combined bill arrive each month — makes it possible to prioritise fixes and hold usage accountable.
Cost discipline as part of adoption, not an afterthought
The other trap is treating AI cost control as a one-off clean-up: review the spend, fix what’s obviously wasteful, move on. That helps once. Then the same sprawl starts again, because the underlying pattern — fast, ungoverned adoption — is usually still there, and new AI tools keep arriving.
The more durable fix is building cost and architecture thinking into how AI gets adopted going forward, so “what will this cost at scale, and does it need to run on the most expensive model available?” is part of the decision from the start, not a question asked eight months later when the invoice has become impossible to ignore.
The commercial case for getting this right
For a tech-focused SMB, AI spend isn’t a novelty line item anymore — it’s rapidly becoming a real input into margins, alongside infrastructure and headcount. Getting the architecture right doesn’t just reduce cost; it usually improves reliability and performance too, because the same discipline that controls spend — matching the right tool to the right task, avoiding redundant work — tends to produce a better-built system regardless.
Getting a proper architectural view of your AI usage, before deciding what to cut or restrict, tends to save more money, preserve more of the value the tools were adopted for, and cause far less disruption than reaching for the off switch.
Ashdown Systems helps UK tech-focused startups and SMBs get their AI adoption — and their AI bills — under control. If any of this sounds familiar, get in touch.







