AI Observability & Cost Monitoring for LLM Apps
Track token usage and cost, and monitor an AI feature in production the same way you would monitor any other service.
Introduction
An AI feature in production has a cost dimension most other backend calls don't: every request has a real, variable dollar cost tied to how many tokens were sent and generated. Lesson 16 covered evaluating quality; this lesson covers the operational side — cost, latency, and reliability once real users are calling your feature.
- Why LLM-powered features need monitoring beyond typical uptime/error tracking.
- How to read and log token usage and cost from a provider response.
- How to think about latency, errors, and budget alerts for an AI feature.
A Real-Life Analogy First
Most home utilities — water, gas, electricity — are metered: the more you use, the more you pay, and a smart meter lets you watch usage climb in real time instead of being surprised by a bill weeks later. A flat-rate subscription like Netflix, by contrast, costs the same no matter how much you use it. Calling an LLM API is a metered utility, not a flat subscription — every request adds to the bill based on exactly how much it used, which is exactly why this lesson exists: without a "meter" of your own, you find out the total only when the bill already arrived.
Why AI Features Need Their Own Monitoring
Variable Cost Per Request
Unlike a typical API call, cost scales directly with input and output length — a single unusually long request can cost far more than average.
Higher, More Variable Latency
Generation time scales with output length, and can vary noticeably run to run, unlike most deterministic backend calls.
Silent Quality Regressions
A provider's model update can change output quality without any error being raised — evaluation tooling (lesson 16) is part of catching this.
Provider Outages
Your feature's uptime is now partly dependent on a third party's uptime, which you don't control.
Tracking Token Usage & Cost
Use case: every provider SDK response includes a usage object reporting exactly how many input and output tokens a call consumed — logging this on every request is the foundation of any cost dashboard, the equivalent of installing your own real-time usage meter instead of waiting for the monthly bill.
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": "Summarize the concept of caching in one sentence."}],)
usage = response.usagecost_per_1k_input = 0.00015 # example pricing, check current provider ratescost_per_1k_output = 0.0006
estimated_cost = ( (usage.prompt_tokens / 1000) * cost_per_1k_input + (usage.completion_tokens / 1000) * cost_per_1k_output)
print(f"Input tokens: {usage.prompt_tokens}, Output tokens: {usage.completion_tokens}")print(f"Estimated cost: ${estimated_cost:.6f}")Click Run to see what this code prints.
Sending usage data to your existing logging or metrics pipeline (alongside a user or feature ID) is what turns "we think AI costs are reasonable" into an actual per-feature cost dashboard you can act on.
Monitoring Latency & Errors
Wrap every provider call with standard timing and error handling, exactly as you would for any external API dependency — measure time-to-first-token separately from total generation time for streaming responses, since the two matter differently for perceived responsiveness in a chat UI.
import time
start = time.monotonic()try: response = client.chat.completions.create( model="gpt-4o-mini", messages=[{"role": "user", "content": "Ping"}], ) latency = time.monotonic() - start print(f"Success in {latency:.2f}s")except Exception as e: latency = time.monotonic() - start print(f"Failed after {latency:.2f}s: {e}") # log to your existing error tracking (Sentry, Datadog, etc.) hereClick Run to see what this code prints.
Setting Budgets and Alerts
Most providers let you set a spending limit or usage alert directly in their dashboard — a critical safety net against a bug (like an infinite agent loop from lesson 11 without a termination condition) silently running up a large bill before anyone notices. This is the direct AI-development equivalent of the spending-cap habit already introduced in lesson 3.
Common Mistakes
- Shipping an AI feature to production with no cost tracking at all, discovering the actual spend only when the bill arrives.
- Treating a provider timeout or rate limit error the same as any other failure, instead of implementing retries with backoff for transient issues.
- Not setting a hard spending cap in the provider dashboard as a last-resort safety net.
Best Practices
- Log token usage, latency, and cost on every request from day one, not after a cost concern comes up.
- Set a provider-side spending alert or hard cap before launching any AI feature, however small.
- Track cost and latency per feature or endpoint, not just as one global number, so you can see exactly what is expensive.
Frequently Asked Questions
Yes — tracing tools like LangSmith typically surface token usage and estimated cost per traced run alongside the execution trace itself.
Tokenizer libraries (like the one shown in lesson 8) can estimate input token count in advance, though output length is only known after generation completes.
For identical or near-identical repeated requests, yes — caching avoids redundant cost and latency, though it requires care around when cached content becomes stale.
The spending cap in lesson 3 is non-negotiable even for a hobby project — it costs nothing to set up and prevents the worst-case scenario. Detailed dashboards and per-feature tracking, on the other hand, are genuinely optional until a project has real, ongoing usage worth optimizing.
Key Takeaways
- AI features carry a variable per-request cost that traditional backend monitoring doesn't typically track — like a metered utility bill, not a flat subscription.
- Every provider response's usage object reports exact token counts for cost calculation.
- Latency should be measured and monitored the same way as any other external dependency.
- Setting a provider-side spending cap is a critical safety net against runaway costs.
Summary
Treating an AI feature's cost, latency, and reliability with the same rigor as any other production dependency is what separates a stable, sustainable AI feature from one that quietly becomes a budget or reliability problem.
- You understand why AI features need dedicated cost and latency monitoring.
- You can read and log token usage and estimated cost from a response.
- You know how to set up spending alerts as a safety net.