The Demo That Lied

Part of the free Generative AI course on LogicWiz, module: Shipping Nova Safely: Observability, Guardrails, Evaluation & Cost.

Episode 46: The Demo That Lied

"It passed every test I gave it," Arjun said. Priya slid her phone across the desk: a subscriber, told the refund window was 30 days, had blown past the real 14-day cutoff. "It didn't crash. That's the problem."


The Demo That Lied

Nova's support desk works. In the demo you type a question, she answers, everyone nods. So Arjun shipped it to LogicWizNews's subscribers — and that is exactly when the trouble started, because a demo and a production system are not the same animal.

Traditional software fails loudly: a crash, a red stack trace, a pager at 2 a.m. You know the instant something broke. AI agents fail silently. Nova hands back a fluent, confident, perfectly-formatted answer that happens to be wrong. Nothing gets paged. The subscriber believes it. You find out weeks later from an angry email.

{{visual:silent-failure}}

That one difference — a wrong answer looks exactly like a right one — is why "it works in the demo" means almost nothing. The rest of this chapter is about the four layers that let you tell a good answer from a confident lie.

Meet Nova's Support Desk

Before we can protect Nova, let's be precise about what she is. Nova isn't one model — she's a small team of agents with a dispatcher out front.

Here is the shape. A subscriber's message first hits a Supervisor — we'll call this step triage. Triage does one job: read the message and decide what kind of question this is. Then it routes to exactly one specialist:

  • KB Agent — answers how-to and policy questions by searching Nova's help articles (that's RAG: retrieve the relevant article, then answer from it).
  • Account Agent — answers questions about one specific subscriber by looking them up in the records ("what plan is SUB-4821 on?").
  • Care Agent — handles upset or sensitive messages with an empathetic, careful reply and no risky guessing.

{{visual:agent-routing-map}}

🧭 Quick refresher: RAG (Retrieval-Augmented Generation) means the agent first fetches relevant text, then answers from that text instead of from memory — the technique you built across Chapters V, VI and X. The KB Agent is a RAG agent.

Notice how simple the flow is: one hop, triage then specialist then answer. That simplicity is exactly why it's fragile — there are more places to go wrong than you'd think.

Where It Can Break

Walk one message through the pipeline and count the failure points. Any single one of these produces a wrong answer that looks perfectly fine:

  • Misroute. Triage sends an account question to the KB Agent. Now Nova searches help articles for a subscriber's balance — and confidently invents one.
  • Retrieval miss. The KB Agent searches, but the vector search returns the wrong article (or nothing), so the model answers from thin air.
  • Context loss. The right article was found, but formatting truncated the key sentence before it reached the model.
  • Hallucination. The right context arrived, but the model ignored it and made up "30 days" anyway.
  • Leak or unsafe reply. The answer is correct but slips in another subscriber's email — or a user tricked Nova into dumping raw records.

Five links in the chain, five ways to fail — and from the outside, every one produces the same thing: a smooth paragraph. You cannot debug what you cannot see, which is where the first layer comes in.

The Four Layers That Make It Real

A production-grade agent is a demo agent wrapped in four operational layers. Each catches a failure the demo never noticed:

{{visual:production-layers}}

  • Observability (Episode 47) — see what happened. End-to-end traces of every step, so when an answer is wrong you can tell whether triage, retrieval, or the model was at fault.
  • Evaluation (Episode 48) — measure whether it's any good. A scored test set, so every prompt tweak is proven better rather than just different.
  • Guardrails (Episode 49) — enforce safety deterministically. Checks on the way in and the way out that a prompt alone can never guarantee.
  • Cost (Episode 50) — afford it at scale. Per-query token tracking, so a runaway bill has a name attached.

We finish with testing and bias (Episode 51), which stitch these layers into a repeatable "is Nova safe to ship?" check.

Build the Triage Step

Enough theory — let's build the first real piece. Triage is the front door: it reads each message and decides which specialist should answer. Get it wrong and every downstream layer ends up protecting the wrong agent — so it's worth getting right. The good news: it's just one small, honest model call. Let's build it up in three ideas.

1. The labels come from Nova's team

Triage doesn't invent categories out of thin air. It has exactly one label per specialist you met earlier in this episode:

  • article → the KB Agent — a how-to or policy question, answered from the help articles
  • account → the Account Agent — something about one specific subscriber
  • care → the Care Agent — an upset or sensitive message

So the label set is simply the menu of things Nova can do. Add a fourth specialist someday and you add a fourth label — no more, no fewer.

2. Triage is a translation into one of those labels

The whole job is a translation: turn an unpredictable sentence — anything a subscriber might type — into exactly one of those three words. A few examples:

  • "How do I cancel my annual subscription?" → article (a how-to, answerable from the help articles)
  • "What plan is subscriber SUB-4821 on?" → account (needs data about one specific subscriber)
  • "I've been charged twice this month and I'm furious!" → care (upset — handle with a human touch)
  • "Do you have a student discount?" → article (still just a policy question)

Pull that off, and the rest of the system is plain, boring code (the word account just means "call the Account Agent"). Here's how the translation works, in three steps:

  1. Classify with the model. Send a short system prompt that names each label in one line and ends with a strict rule: "reply with ONLY the label word." That rule is the trick — a one-word answer is something code can act on directly, instead of a sentence you'd have to parse.
  2. Make it repeatable with temperature=0. Temperature is the model's randomness dial (from Chapter III). Classification wants the opposite of creativity — the same message should always get the same label — so we turn the randomness off.
  3. Never trust the raw reply. Models are probabilistic; once in a while one returns something off-menu ("billing", or nothing at all). So we check: is the reply one of our three labels? If yes, use it; if not, fall back to a safe default (article) rather than crash or route nowhere.

3. Routing is then a one-line lookup

With a clean label in hand, routing is trivial: a plain dictionary maps it to the specialist — {"article": "KB Agent", "account": "Account Agent", "care": "Care Agent"}. No AI, just a lookup. That's the pattern you'll see all chapter long: let the model do the fuzzy part (understanding the message), then hand off to deterministic code for anything that has to be reliable.

Now walk the real code, line by line — it's exactly these three ideas:

{{visual:triage-walkthrough}}

You'll write this classifier — and the routing that follows — yourself in this lesson's lab.

💡 Why a fixed label set and a safe default? Everything downstream is deterministic code that switches on the label. If triage ever returns a surprise word, the default keeps Nova answering instead of crashing — a tiny but real production habit.

What Nova Learns Next

Nova can route a message now — but when she gets an answer wrong, Arjun still has no idea which step failed. In Episode 47 we build the first production layer: observability — the traces that turn a silent failure into a story you can actually read.