Measuring Instead of Guessing

Part of the free Generative AI course on LogicWiz, module: Shipping Nova Safely: Observability, Guardrails, Evaluation & Cost.

Episode 48: Measuring Instead of Guessing

"I fixed the refund answer," Arjun announced. Priya checked another ticket: now the billing answer was broken. "You didn't fix it," she said. "You moved it."


The Trap of Prompt Gambling

Arjun's debugging loop looked productive: spot a wrong answer, tweak the prompt, eyeball a couple of questions, ship. But a prompt change ripples everywhere. Tighten it to fix refunds and you might quietly wreck billing. Without a way to measure the whole system, every "fix" is a gamble — you're trading one visible bug for an invisible one.

The cure is evaluation: run Nova against a fixed set of questions with known answers, and get a score. Now a change isn't "it feels better" — it's "correctness went 0.62 → 0.78." Evaluation is just testing, for a system whose output isn't exact — you can't assert answer == "14 days" on free text, so you score it instead.

Evaluation Is a Stack

Here's the first surprise: there's no single "is Nova good?" number. A wrong answer fails at some specific step, so evaluation is a stack of layers, each checking one step of the pipeline:

{{visual:eval-layers}}

  • Routing — did triage pick the right specialist? (A plain match: 1.0 or 0.0.)
  • Retrieval — did we fetch the right article, and how high did it rank? (Measured with MRR, Mean Reciprocal Rank: if the correct doc is 1st you score 1.0, 2nd → 0.5, 3rd → 0.33. The higher the right answer ranks, the better.)
  • Faithfulness — is the answer grounded in what we retrieved?
  • Correctness — does it match the true answer?
  • Quality — is it well-written, empathetic, safe?

Layering matters because it tells you where to fix. A low retrieval score says "fix the search," not "fix the prompt." One number would hide that.

Faithful Isn't Correct

The two middle layers look similar and are constantly confused — get them straight and half of evaluation clicks:

{{visual:faithfulness-vs-correctness}}

  • Faithfulness asks: does the answer stay inside the retrieved context? An answer can be perfectly faithful and still wrong — if the context itself was the wrong article. (This is exactly the refund bug from Episode 47.)
  • Correctness asks: does the answer match the real, ground-truth answer, regardless of context?

You need both. Faithfulness catches hallucinations; correctness catches retrieving-the-wrong-thing. Check only one and a whole class of silent failures walks right through.

The LLM as Judge

Some layers are easy to score with plain code: routing is a string match, and correctness can often check for a keyword or a number. But how do you score faithfulness — "is every claim supported?" — programmatically? You can't. So you use a second model as a judge.

"Isn't AI grading AI circular?" It's the first thing everyone asks, and the answer is no: grading is far easier than answering. The judge doesn't need to know Nova's refund policy — it only has to check whether the answer stays inside the context it was handed. Verifying is easier than generating.

Here's the judge you'll build in the lab — a strict rubric, forced JSON, one score:

{{visual:llm-judge-walkthrough}}

Judges aren't magic, so production teams harden them: temperature 0 and a sharp rubric for consistency, averaging over a few runs to smooth noise, a cross-family judge (have one model answer and a different model grade, so a model can't rate its own quirks kindly), and guarding the parse so a malformed reply defaults to a neutral score instead of crashing.

💡 Rule of thumb: reach for the LLM judge only where plain code can't reason — faithfulness, tone, empathy. For routing and exact facts, hardcoded logic is cheaper, faster, and perfectly repeatable.

Hill-Climbing: One Lever at a Time

Now you can measure, so you can improve — deliberately. First you need a dataset: a handful of realistic questions, each with its expected intent and answer. (Even 10–15 good examples beat eyeballing.) Then you improve by hill-climbing:

{{visual:hill-climbing}}

The golden rule is the whole visual: change exactly one variable per experiment. Bump top_k, re-run the eval, keep it only if the score rose; then change the prompt, re-run, keep or revert. Change two things at once and a win tells you nothing — you don't know which one helped. Each kept change becomes the new baseline, and the score climbs a step at a time.

The same eval that powers hill-climbing becomes a quality gate in CI: block any change that drops correctness below a threshold, and Nova can never silently regress on the way to production.

What Nova Learns Next

Nova is now measurable — Arjun can prove a change helps instead of hoping. But evaluation scores typical questions. It won't stop a malicious one: a user who tries to make Nova leak a subscriber's data or ignore her rules. In Episode 49 we build the third layer: guardrails — deterministic checks that hold even when a prompt is under attack.