top of page

Elevated Magazines - Premium Lifestyle Content

From the superyachts making waves at Monaco to the estates redefining luxury living in Palm Beach, the automotive debuts turning heads in Geneva, and the artists commanding record prices at auction — Elevated Magazines captures the luxury lifestyle stories, brands, and cultural moments that have the world's most discerning audiences talking right now.

Your LLM Bill Is Not a Model Pricing Problem

5 days ago
5 min read

When inference costs come up in an engineering review, the conversation almost always starts at the wrong layer. Someone compares per-token prices across providers, someone else suggests a smaller model, and the meeting ends with a plan to renegotiate.

Per-token price is usually the least controllable variable in the stack. The controllable ones sit above and below it, and they can move your bill by a factor of several without changing provider at all.

Layer one: how many tokens you send

The largest single cost lever in most production systems is prompt size, and the largest contributor to prompt size is context that nobody has pruned since the feature shipped.

Typical patterns worth auditing: a system prompt that has accumulated instructions over eighteen months and is now two thousand tokens on every call; retrieved context set to a fixed number of chunks regardless of query; entire conversation histories resent rather than summarised; few-shot examples retained long after the model stopped needing them.

None of this is exotic. It is the software equivalent of leaving debug logging on in production, and it is extremely common because token counts are invisible unless someone deliberately looks.

The first exercise for any team with a surprising bill is to log token counts per request type and sort descending. The answer is usually one or two endpoints, and the fix is usually an afternoon.

Layer two: the serving engine

Below the application sits the inference engine, and for self-hosted deployments this is where the largest throughput differences live.

The mechanics that matter are prefix caching and scheduling. When many requests share a common prefix — the same system prompt, the same document, the same conversation head — an engine that recognises and reuses that computation serves dramatically more traffic on the same hardware than one that recomputes it. Different engines implement this differently, and the implementation interacts with your traffic shape rather than being universally better.

This is why benchmark numbers from someone else's workload are close to useless for capacity planning. Edgewisely's comparison of two major inference engines and how their caching strategies suit different traffic patterns makes the point that the architectural difference, not the headline throughput figure, is what determines which one suits you. A workload with long shared prefixes and short completions behaves nothing like one with short prompts and long generations.

Measure with a replay of your own traffic. Anything else is a guess dressed up as a number.

Layer three: routing and the gateway

Between your application and the models sits some form of gateway, and the decision about what goes there has cost implications people rarely model.

Hosted aggregators are convenient: one integration, many models, instant failover, no infrastructure. They take a percentage. Self-hosted gateways are free of that percentage and cost you operational ownership — deployment, upgrades, an on-call rotation, and the failure mode where your own proxy is the outage.

The percentage sounds small until your spend is large, at which point it funds an engineer. Edgewisely's breakdown of a self-hosted gateway against a hosted aggregator, including the real fee structure makes the crossover point clear, and the honest answer is that the right choice changes as you scale rather than being a permanent architectural preference.

The more valuable thing a gateway gives you, at either scale, is the ability to route by task. Most production systems have a handful of request types that genuinely need a frontier model and a long tail that a cheaper one handles identically. Teams that route uniformly are paying frontier prices for classification and formatting. This single change typically delivers more saving than any provider negotiation, and it requires an evaluation harness to do safely — which is the real reason most teams have not done it.

Layer four: the hardware, if you own it

For self-hosted deployments, the hardware question is usually framed as a generational upgrade decision, and the marketing is unusually misleading here because headline comparisons are frequently made at different numerical precisions.

A figure quoting a large multiple of improvement often compares a new part at a lower precision against an older part at a higher one. Matched-precision comparisons produce considerably more modest numbers — Edgewisely's benchmark-based comparison of two recent accelerator generations at matched precision finds the real advantage well below the headline multiple, which changes the payback calculation on an upgrade substantially.

The related trap is buying capacity before you have fixed layers one through three. An engine change and a prompt audit routinely deliver more effective capacity than a hardware generation, at zero capital cost. Hardware is the correct answer only after software efficiency has been exhausted, and almost nobody exhausts it first because hardware is procurable and prompt hygiene is not.

The layer everyone forgets: caching your own results

There is a fifth lever that sits outside the stack entirely, and it is the cheapest of all: not making the call.

A surprising share of production traffic is repetitive. The same questions about the same documents, the same classifications of near-identical inputs, the same summaries regenerated because nobody stored the last one. A semantic cache in front of the model — returning a stored answer when a new request is sufficiently similar to a previous one — eliminates that traffic entirely rather than making it cheaper.

The engineering is not trivial. You need a similarity threshold that is tight enough to avoid returning a subtly wrong answer, an invalidation strategy for when the underlying data changes, and a way to exclude anything personalised or time-sensitive. Set the threshold badly and you will serve a confidently wrong cached response, which is worse than the cost you saved.

Done carefully, though, hit rates in the tens of percent are common for document-oriented workloads, and every hit is a request that costs nothing, adds no latency and consumes no capacity. Teams reach for this last, after they have optimised everything about how they make the call, having never asked whether the call needed to happen.

The same reasoning applies upstream: batching requests that do not need to be synchronous, and deferring work that nobody is waiting for, both convert expensive interactive capacity into cheap background capacity.

The audit, in order

If you have a cost problem, work the layers in this sequence, because each one changes the numbers for the ones below.

Log tokens by endpoint and cut the obvious waste. Introduce task-based routing with an evaluation harness so you can prove the cheaper model is adequate. Replay your real traffic against candidate serving engines and choose on your own numbers. Only then price hardware.

Teams that run this sequence typically find substantial reductions without changing provider or model quality. Teams that start at provider negotiation get a small discount on a badly shaped bill, congratulate themselves, and repeat the exercise next year.

The per-token price is the one number everybody looks at and the one nobody controls. Everything you actually control is somewhere else.


Perrelet Casino Royale
Northrop & Johnson Yachts for Charter
Nuvolari Lenard
bottom of page