Seeing Inside Nova
Part of the free Generative AI course on LogicWiz, module: Shipping Nova Safely: Observability, Guardrails, Evaluation & Cost.
Episode 47: Seeing Inside Nova
The refund answer was wrong again. "Which step broke?" Priya asked. Arjun stared at the single paragraph Nova returned. "I... have no idea. I can see what she said — not what she did."
Seeing Inside Nova
Last episode ended with a silent failure: Nova told a subscriber the refund window was 30 days when the real answer is 14. From the outside it's one smooth paragraph — no error, no clue. To fix it, Arjun needs to see inside the request: every step Nova took, in order, with what went in and what came out.
Think about how you debug a late package. You don't just know "it's late" — you open the tracking history: picked up 9:02 → left the warehouse 11:40 → arrived in the wrong city 4:15 → …. Each step is stamped with a time and a place, so you can point at the exact one that went wrong. A trace is that tracking history, but for one request moving through Nova: a timed record of every step, stitched together into a tree.
Play one and watch it turn the silent failure loud:
{{visual:trace-waterfall}}
The trace localizes the bug instantly: retrieval fetched the wrong article, and the model answered faithfully — from the wrong source. Without the trace, "the answer is wrong" could have been triage, retrieval, formatting, or the model. With it, you point straight at the failing step.
Logging, Monitoring, Observability
"Isn't that just logging?" It's the natural question — these three words get thrown around as synonyms, but they work at completely different scales, and mixing them up is why teams end up staring at a dashboard that can't answer the question they actually have. These aren't AI ideas at all — they come straight from running normal web software. Picture a plain online store, with a few services behind it (web server → cart → payment → inventory):
- Logging is the raw server log — every request and error on its own line, each time-stamped:
GET /cart 200,POST /pay 500. Precise, but a firehose of separate lines. - Monitoring is the ops dashboard on the wall — requests per second, error rate, CPU,
P99latency — with an alarm when a number leaves its safe range. It watches the whole system's health at a glance. - Observability is picking one shopper whose checkout failed and following their request across every service — web → cart → payment → inventory — as one connected story, so you can explain exactly why their order broke.
{{visual:logging-vs-observability}}
Nova is the same kind of system — a request hopping through triage, retrieval, and the model instead of cart and payment — so the three tools map straight across. The clearest way to tell them apart is the scale each works at. Here they are side by side:
| Logging | Monitoring | Observability | |
|---|---|---|---|
| Scale | one event | the whole system | one request |
| Answers | "what happened at 10:04:02?" | "is the system healthy right now?" | "why did THIS answer come out wrong?" |
| Looks like | a flat list of log lines | dashboards + alerts on averages | a connected trace of every step |
| Blind spot | events aren't linked into a story | an average hides one bad answer | costly to keep 100%, so you sample |
Now the same three in a bit more depth:
Logging — the micro scale (individual events). A flat, append-only list of things that happened: triage started, search returned 4 docs, response sent. Every line is true and precise. Its limit is that the lines aren't tied together — you can see that a search ran at 10:04:02, but not which request it belonged to or what happened right before and after it. Logging answers "what happened at this instant?"
Monitoring — the macro scale (aggregate health). Numbers rolled up across all requests: average latency, error rate, requests per minute, and percentiles like P99 (the latency the slowest 1% of users actually feel). You watch these on dashboards and wire alerts to them — "page me if the error rate crosses 2%." Its limit: it tells you that something is wrong across the system, never why one specific answer was wrong — the averages can look perfectly healthy while an individual subscriber gets nonsense. Monitoring answers "is the system healthy overall?"
Observability — the relational scale (one request, end to end). Every step of a single request stitched into a trace, each step carrying its inputs, outputs, latency and tokens, and which step called which. It's the only view that reconstructs the story of one request — exactly the trace you played above. Observability answers "why did THIS answer come out the way it did?"
Here's the part people miss: they aren't rivals — you run all three, and they hand off to each other. Walk through a real incident:
- Monitoring fires an alert:
P99latency just doubled. → You know something is wrong. - Observability: you open a slow request's trace and see the
retrievespan is taking 5× longer than usual. → You know where. - Logging: you read the raw log lines for that span and find the vector database timed out and retried. → You know exactly what.
Monitoring caught it, observability localized it, logging confirmed the root cause. Drop the middle layer and step 1 never reaches step 3 — you'd know the system is slow but never which step or why. That gap, for a system that fails silently, is the whole reason this layer exists.
What a Trace Captures
A trace is a tree of spans — one span per step. It's a tree because steps nest inside each other: the kb_agent span contains a retrieve span and an llm.answer span, the same way a function contains the calls it makes. For the refund request, the tree looked like this:
triage.classify— 0.4skb_agent— 1.5sretrieve— 0.3s (returned the wrong article — the bug)llm.answer— 1.2s
That tree tells you the shape and timing. But the real power is inside each span. So let's zoom into one — here is everything the llm.answer span actually recorded on that request:
- step:
llm.answer - parent:
kb_agent(where it sits in the tree) - input: the Billing FAQ text + "How long is the refund window?"
- output: "You have 30 days to request a refund."
- latency: 1.2s
- tokens: 512 in, 18 out
Read that one record and the whole silent failure cracks open. The input shows the model was handed the Billing FAQ — so the instant you also notice retrieve fetched the wrong doc, the case is closed: the model didn't hallucinate, it answered faithfully from bad context. You didn't guess that; you read it straight off the span.
Every span carries this same shape, and its fields are exactly the four things you can't debug without:
- Inputs and outputs — what went in and came out of each step. This is how you saw that
llm.answerwas fed the Billing FAQ instead of the Refunds article. - Latency — how long each step took (those waterfall bars — also perfect for catching the slow step, not just the wrong one; here
llm.answer's 1.2s dominates). - Tokens — how many tokens each model call used (
retrieveused 0 — it's a database lookup, not a model call;llm.answerused 530). This is also where cost tracking begins — Episode 50. - Hierarchy — which step called which. Because
retrieveandllm.answersit underkb_agent, a glance tells you the failure lives inside the KB Agent, not in triage.
So how do you actually get traces? Two ways, depending on your stack:
- The free switch. If you build on the LangChain / LangGraph family, tracing is basically a setting. Flip a couple of environment variables and a tool like LangSmith records every run automatically — no code changes at all.
- Manual instrumentation. If your code isn't in that ecosystem, you add it yourself: wrap your model client so every call is captured, and decorate the functions you want traced. A little more effort, same result — the same tree of spans.
🧭 On tools: LangSmith is the popular hosted option, and its free developer tier is plenty for building and testing. But in this lesson you'll build the mechanism yourself in a few lines — because once you've built a tracer, no hosted tool is a mystery again.
Build a Tracer
A tracer sounds heavy; it isn't. Strip it down and it's just two pieces:
- A list to collect spans — literally
TRACE = []. - A "span" you wrap around any step — it starts a timer, lets that step record its output, and when the step finishes it stamps the elapsed time and drops the record into the list.
The clean way to run code before and after a block in Python is a context manager — the thing you use with a with statement. So writing with span("retrieve"): times whatever runs inside it and files the span automatically, with no bookkeeping cluttering your actual logic. Here's the shape you'll build:
{{visual:tracer-walkthrough}}
In the lab you'll write that span context manager, then wrap Nova's real steps in it and read the trace back — the exact move that turns "the answer is wrong" into "the retrieve step is wrong."
Tracing at Scale
In development you trace everything — volume is tiny and you want every detail. In production that's wasteful: at, say, 200,000 requests a day, storing every trace is a mountain nobody reads (and a real bill). So you sample — keep only a fraction:
{{visual:sampling}}
Two habits make sampling safe:
- Keep a small random slice (often 10–20%) for the everyday health picture.
- Always keep 100% of errored traces (conditional sampling). A random 10% might miss the one failure you needed — but if you save every trace that hit an error, you never lose the ones that matter.
On top of the traces sit monitoring dashboards — token spend, latency percentiles, how often each agent is called — with alerts that page you when latency spikes or cost crosses a line. Traces tell you why one request failed; dashboards tell you when the whole system starts to drift.
What Nova Learns Next
Arjun can now see exactly what Nova did on any request. But seeing one failure isn't the same as knowing whether a fix helps — change a prompt to patch this bug and you might quietly break ten other answers. In Episode 48 we build the second layer: evaluation — measuring Nova's quality across a whole test set, so every change is proven better instead of just different.