The other day in a review meeting: a client proudly showed us his agent landscape — and then the cloud bill. The enthusiasm in the room cooled noticeably. It was the moment a technology topic became an economics topic.

Small, fast, surprisingly good

Small language models — compact language models with a fraction of the parameters of large frontier models — have reached a remarkable level of maturity in recent months. For clearly defined tasks such as classification, extraction, routing or summarisation, they deliver results that are hardly distinguishable from those of the big ones — at a tenth to a thirtieth of the cost and with considerably lower energy consumption. The rule of thumb that has become established with us: the big model thinks, the small model works.

FinOps for agents: the new mandatory discipline

As soon as agents run around the clock, every unnecessary model call becomes a lever. AI FinOps — the cost control of AI workloads — means in practice: routing tasks by difficulty instead of sending everything to the most expensive model, caching intermediate results, keeping contexts lean and measuring per process what a completed task actually costs. Only once the price per case is known can scaling be discussed seriously.

What this means for the mid-market

The good news: this development plays into the hands of smaller companies. Those who don't have the budget for millions of frontier-model calls often do better with a smart mix of small, specialised models and a top model switched in selectively — faster, cheaper and with less dependence on a single provider. Efficiency is no longer a compromise, it is the strategy.

If you want to know what your AI processes may cost per case — and how to get there — we are happy to work it out together.

Datenschutz-Einstellungen