LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 2323 min read

AI Observability & Cost Monitoring for LLM Apps

Track token usage and cost, and monitor an AI feature in production the same way you would monitor any other service.

Introduction

An AI feature in production has a cost dimension most other backend calls don't: every request has a real, variable dollar cost tied to how many tokens were sent and generated. Lesson 16 covered evaluating quality; this lesson covers the operational side — cost, latency, and reliability once real users are calling your feature.

What You Will Learn
  • Why LLM-powered features need monitoring beyond typical uptime/error tracking.
  • How to read and log token usage and cost from a provider response.
  • How to think about latency, errors, and budget alerts for an AI feature.

A Real-Life Analogy First

Most home utilities — water, gas, electricity — are metered: the more you use, the more you pay, and a smart meter lets you watch usage climb in real time instead of being surprised by a bill weeks later. A flat-rate subscription like Netflix, by contrast, costs the same no matter how much you use it. Calling an LLM API is a metered utility, not a flat subscription — every request adds to the bill based on exactly how much it used, which is exactly why this lesson exists: without a "meter" of your own, you find out the total only when the bill already arrived.

Why AI Features Need Their Own Monitoring

Variable Cost Per Request

Unlike a typical API call, cost scales directly with input and output length — a single unusually long request can cost far more than average.

Higher, More Variable Latency

Generation time scales with output length, and can vary noticeably run to run, unlike most deterministic backend calls.

Silent Quality Regressions

A provider's model update can change output quality without any error being raised — evaluation tooling (lesson 16) is part of catching this.

Provider Outages

Your feature's uptime is now partly dependent on a third party's uptime, which you don't control.

Tracking Token Usage & Cost

Use case: every provider SDK response includes a usage object reporting exactly how many input and output tokens a call consumed — logging this on every request is the foundation of any cost dashboard, the equivalent of installing your own real-time usage meter instead of waiting for the monthly bill.

from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize the concept of caching in one sentence."}],
)
usage = response.usage
cost_per_1k_input = 0.00015 # example pricing, check current provider rates
cost_per_1k_output = 0.0006
estimated_cost = (
(usage.prompt_tokens / 1000) * cost_per_1k_input
+ (usage.completion_tokens / 1000) * cost_per_1k_output
)
print(f"Input tokens: {usage.prompt_tokens}, Output tokens: {usage.completion_tokens}")
print(f"Estimated cost: ${estimated_cost:.6f}")
Terminal Output

Click Run to see what this code prints.

Log This on Every Call, Not Just in Testing

Sending usage data to your existing logging or metrics pipeline (alongside a user or feature ID) is what turns "we think AI costs are reasonable" into an actual per-feature cost dashboard you can act on.

Monitoring Latency & Errors

Wrap every provider call with standard timing and error handling, exactly as you would for any external API dependency — measure time-to-first-token separately from total generation time for streaming responses, since the two matter differently for perceived responsiveness in a chat UI.

import time
start = time.monotonic()
try:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Ping"}],
)
latency = time.monotonic() - start
print(f"Success in {latency:.2f}s")
except Exception as e:
latency = time.monotonic() - start
print(f"Failed after {latency:.2f}s: {e}")
# log to your existing error tracking (Sentry, Datadog, etc.) here
Terminal Output

Click Run to see what this code prints.

Setting Budgets and Alerts

Most providers let you set a spending limit or usage alert directly in their dashboard — a critical safety net against a bug (like an infinite agent loop from lesson 11 without a termination condition) silently running up a large bill before anyone notices. This is the direct AI-development equivalent of the spending-cap habit already introduced in lesson 3.

Common Mistakes

Avoid These Mistakes
  • Shipping an AI feature to production with no cost tracking at all, discovering the actual spend only when the bill arrives.
  • Treating a provider timeout or rate limit error the same as any other failure, instead of implementing retries with backoff for transient issues.
  • Not setting a hard spending cap in the provider dashboard as a last-resort safety net.

Best Practices

  • Log token usage, latency, and cost on every request from day one, not after a cost concern comes up.
  • Set a provider-side spending alert or hard cap before launching any AI feature, however small.
  • Track cost and latency per feature or endpoint, not just as one global number, so you can see exactly what is expensive.

Frequently Asked Questions

Yes — tracing tools like LangSmith typically surface token usage and estimated cost per traced run alongside the execution trace itself.

Tokenizer libraries (like the one shown in lesson 8) can estimate input token count in advance, though output length is only known after generation completes.

For identical or near-identical repeated requests, yes — caching avoids redundant cost and latency, though it requires care around when cached content becomes stale.

The spending cap in lesson 3 is non-negotiable even for a hobby project — it costs nothing to set up and prevents the worst-case scenario. Detailed dashboards and per-feature tracking, on the other hand, are genuinely optional until a project has real, ongoing usage worth optimizing.

Key Takeaways

  • AI features carry a variable per-request cost that traditional backend monitoring doesn't typically track — like a metered utility bill, not a flat subscription.
  • Every provider response's usage object reports exact token counts for cost calculation.
  • Latency should be measured and monitored the same way as any other external dependency.
  • Setting a provider-side spending cap is a critical safety net against runaway costs.

Summary

Treating an AI feature's cost, latency, and reliability with the same rigor as any other production dependency is what separates a stable, sustainable AI feature from one that quietly becomes a budget or reliability problem.

Lesson 23 Completed
  • You understand why AI features need dedicated cost and latency monitoring.
  • You can read and log token usage and estimated cost from a response.
  • You know how to set up spending alerts as a safety net.
Next Lesson →

Real-World Project: Building an AI Tooling Stack