SDP Clouds
AI·4 min read

LLM Cost and Latency: The Tradeoffs Behind the Pricing Page

Bigger model, better answer, four times the bill and triple the p95. How to reason about model size, caching, and routing so unit economics survive real traffic.


The demo used the biggest model and the answer was excellent. Production used the biggest model and the bill was discussed in a meeting with finance in it.

Neither conversation was wrong. They were just answering different questions: "can this do the task at all?" versus "can this do the task two million times a month at a price the product can carry?" Most teams ship the first answer and discover the second one in month two.

The two numbers nobody puts on the pricing page

Every model is sold on quality. What you actually buy is a pair: cost per successful outcome and p95 latency at your traffic shape.

Cost per outcome is the honest unit, not cost per token. A model that is 4× cheaper per token but needs two retries to get a usable answer is more expensive than the one you were trying to replace. This is exactly why evaluations are a cost tool and not just a quality tool — without a pass rate you cannot compute the number finance is asking for.

Latency compounds differently. Users tolerate a slow answer that is right; they do not tolerate a fast one that is wrong, because the retry lands on top of the original wait. Your p95 is a product feature whether or not you have declared it one.

Where the latency actually goes

Before optimising anything, find out what you are optimising. The span usually breaks down as:

  • Queue and cold start — worst on serverless endpoints, variable by time of day.
  • Time to first token — dominated by prompt length and prefill. A 12,000-token system prompt you paste in every request is paid on every request.
  • Generation — the part people imagine, and often the smallest of the three.
  • Your own post-processing — validators, retries, and a JSON parse that fails and re-runs the whole call.

Instrument the fourth one first. Teams routinely discover that a schema-validation retry is doubling their spend while they are busy benchmarking model choice.

The router, not the bigger model

The cheapest optimisation available is sending each request to the smallest model that can handle it, and most requests are not the hard ones.

A rough split that holds more often than it should:

Request shareModel choice
~60%format, classify, extract — small model, or no model at all
~30%genuine reasoning, but bounded — mid-tier model
~10%the cases that made you want this feature — largest model

The first row is the one that pays for the other two, and it is frequently a rules engine that somebody replaced with a model in a hurry. If a regex does it, let the regex do it.

Caching: the discount nobody applies

If your system prompt or your few-shot examples are stable, you are re-sending and re-paying for them on every call. Prompt caching changes the arithmetic substantially — often the single largest line-item reduction available, and it requires no change to your output quality.

Cache what is genuinely stable; do not cache what changes per user, or you will trade a discount for correctness bugs that only appear under concurrency.

What I would measure first

Before changing models or providers, get three numbers for the current setup: cost per successful outcome (spend ÷ tasks that did not need a human), p95 end to end, and retry rate. Most cost surprises turn out to be the third number wearing a costume.

Then make changes one at a time, with the same three numbers after each. The pattern in our own pipeline was the same as everyone else's: the wins came from routing and caching, not from swapping the flagship model for a different flagship model.

Summary

Price the outcome, not the token — a cheap model that retries is not cheap. Break latency into queue, prefill, generation, and your own retries before touching anything, because the retry is often the bug. Route the easy sixty percent to something small, cache what is stable, and measure cost-per-successful-outcome, p95, and retry rate as the three numbers that actually decide whether the feature survives contact with finance.

#llm#ai#cost-optimization#latency#ai-engineering

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles