The Bill Nobody Watched
Part of the free Generative AI course on LogicWiz, module: Shipping Nova Safely: Observability, Guardrails, Evaluation & Cost.
Episode 50: The Bill Nobody Watched
The invoice landed and Arjun blinked. Last month's model spend had tripled. "Which feature did that?" Priya asked. He had absolutely no idea.
The Bill Nobody Watched
Nova is observable, measured, and safe — and quietly expensive. The trap is that each model call feels like a rounding error, so nobody watches the total. But a single subscriber message is never a single call:
{{visual:cost-multiplier}}
Every message runs triage (one call) plus a specialist (another call), often with a guard classifier on top — three calls, tens of thousands of times a day. The "cheap" model is cheap per call; multiply by the multi-agent fan-out and the daily volume, and the bill is real. Worse, without tracking you can't answer the only question that matters when the bill spikes: which user, feature, or query caused it?
Why Output Costs More
Here's the fact that surprises everyone: output tokens cost several times more than input tokens — typically about 4×. Reading your prompt is cheap; generating a reply is where the money goes.
{{visual:token-economics}}
That flips your intuition about optimization. A long prompt is not the enemy — a long answer is. The single cheapest win is often just instructing Nova to answer concisely: fewer output tokens, a real cost cut, and usually no loss in quality (subscribers prefer short answers anyway).
Count Before You Spend
You can't optimize what you don't measure, and you don't need to spend a cent to start. The models tokenize text with a library called tiktoken — the same tokenizer they use internally — and you can run it locally to count tokens before any API call:
{{visual:cost-walkthrough}}
With a token count and a price table, every query gets a real dollar figure. Attach that figure to each request's intent (article / account / care) and you can finally see where the money goes — maybe the care agent's long, careful replies cost 5× a quick account lookup. That per-intent breakdown is exactly what you'll build in the lab.
💡 tiktoken gives you input cost up front. For the exact output cost you read the token usage the API returns after the call (the same tokens your tracer already records — Episode 47). Estimate with tiktoken, confirm with the real usage.
Cutting the Bill Without Cutting Quality
Once you can measure, optimization is a set of deliberate trades — each measured, one variable at a time:
- Answer concisely. The output-token lever from earlier — biggest win for the least effort.
- Route by difficulty. Send easy queries to the small model and reserve the big, pricey model for the hard ones. Most traffic is easy.
- Trim the context. Retrieve fewer, better chunks (a smaller
top_k) so each call carries less input. - Cache repeats. Identical or near-identical questions can serve a stored answer instead of paying for the model again.
The goal is never "cheapest" or "best" — it's the knee of the curve: most of the quality for a fraction of the cost.
{{visual:cost-pareto}}
A Cost Dashboard
Measurement only helps if someone sees it. In production the per-query costs feed a dashboard — spend by day, by intent, by model — with alerts: page someone if the daily budget is exceeded, or if a single query costs more than a threshold (a runaway loop, usually). The habit is the same as observability: track the number, set a line, and get told automatically when you cross it — before the invoice does the telling.
What Nova Learns Next
Nova is now observable, measurable, safe, and affordable — every production layer is in place. One question remains before Arjun can trust her with real subscribers at scale: how do you keep her that way, and make sure she treats every subscriber fairly? In Episode 51 we close the chapter with testing and bias — the checks that make shipping Nova a decision you can defend.