LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1623 min read

Prompt Engineering & Evaluation Tools

Learn LangSmith for tracing LLM calls and Promptfoo for automated prompt evaluation, with a working example of each.

Introduction

Traditional software has deterministic tests: same input, same output, every time. LLM output is not deterministic in the same way — the same prompt can produce a slightly different answer twice, and a prompt change that improves one case can silently break another. Evaluation tooling exists to catch that.

What You Will Learn
  • Why manual "it worked when I tried it" testing is not sufficient for LLM applications.
  • How to trace a chain's execution with LangSmith.
  • How to run automated, repeatable prompt evaluations with Promptfoo.

A Real-Life Analogy First

Think about a restaurant kitchen. A chef tasting one plate before it goes out is quality control for a single dish, in the moment. A health inspector who visits regularly, checks records, and re-tests standard procedures is a completely different kind of check — one that catches problems a single taste test never could, like a recipe that's only sometimes prepared incorrectly, or ingredients that quietly change quality over time. LLM applications need both kinds of check: a quick "did this one response look right" (a manual spot check) and a systematic, repeatable inspection (evaluation tooling).

Why "It Worked Once" Isn't Enough

A RAG pipeline (lesson 15) or agent (lesson 11) has many moving parts — a retrieval step, a prompt template, a model call — and a bug can hide in any one of them. Without visibility into each individual step, "the final answer looked wrong" gives you almost no information about where in the pipeline the problem actually happened — the same way "the meal didn't taste right" tells a health inspector nothing about which specific step in the kitchen went wrong.

LangSmith: Tracing & Observability

Use case: LangSmith (from the LangChain team) records a full trace of every step in a chain's execution — each prompt sent, each model response, each retrieved document — viewable in a dashboard, so you can see exactly what happened on any given run, not just the final output.

pip install langsmith

LangSmith in Action

# .env
LANGCHAIN_TRACING_V2=true
LANGCHAIN_API_KEY=your-langsmith-key
LANGCHAIN_PROJECT=support-chatbot

With those three environment variables set, any LangChain chain (like the one built in lesson 9) automatically sends a full trace of each run to LangSmith — no code changes needed beyond loading the environment variables, since LangChain checks for them at startup.

What You See in the Dashboard

Click Run to see what this code prints.

Promptfoo: Automated Prompt Testing

Use case: Promptfoo runs a prompt against a defined set of test cases and expected criteria automatically, similar in spirit to a unit test suite — so a prompt or model change can be checked against every known case before shipping, not just the one example you happened to try by hand.

npm install -g promptfoo

Promptfoo in Action

# promptfooconfig.yaml
prompts:
- "Answer the support question concisely: {{question}}"
providers:
- openai:gpt-4o-mini
tests:
- vars:
question: "How long do refunds take?"
assert:
- type: contains
value: "5-7 business days"
- vars:
question: "Do you ship internationally?"
assert:
- type: llm-rubric
value: "The answer should not guess -- it should say this information isn't available."
promptfoo eval
Terminal Output

Click Run to see what this code prints.

Common Mistakes

Avoid These Mistakes
  • Shipping a prompt change without re-running any evaluation, relying purely on a single manual spot check.
  • Writing only "happy path" test cases and never testing edge cases like missing context or ambiguous questions.
  • Treating tracing and evaluation as optional "nice to haves" added only after a production incident.

Best Practices

  • Add tracing (LangSmith or an equivalent) from the very first prototype, not after something breaks in production.
  • Build a small evaluation set (even 10-15 cases) covering both expected answers and known edge cases.
  • Re-run your evaluation suite on every prompt or model change, the same way you would run tests on every code change.

Frequently Asked Questions

Not necessarily — they solve related but distinct problems: LangSmith focuses on tracing what actually happened, Promptfoo focuses on repeatable pass/fail testing. Many teams use both together.

LangSmith integrates most seamlessly with LangChain, but also offers a standalone SDK for tracing raw provider SDK calls directly.

Yes — it supports most major providers covered in lessons 5 and 6, plus custom provider configuration for anything else.

For a five-minute experiment, yes, skip it. The moment a project has real users, or you plan to keep changing the prompt over time, even a tiny evaluation set (2-3 test cases in a Promptfoo config) pays for itself the first time it catches a change that quietly broke something you weren't looking at.

Key Takeaways

  • LLM output's non-determinism makes manual, one-off testing unreliable.
  • LangSmith traces every step of a chain's execution for debugging and observability.
  • Promptfoo runs automated, repeatable evaluations against defined test cases.
  • Evaluation should be part of the development loop from the start, not bolted on after an incident.

Summary

Tracing and evaluation tooling turn "it seemed to work" into something you can actually verify and re-check on every change — a discipline that separates a demo from a production-ready AI feature.

Lesson 16 Completed
  • You understand why LLM applications need dedicated evaluation tooling.
  • You can trace a chain's execution with LangSmith.
  • You can run automated prompt tests with Promptfoo.
Next Lesson →

Fine-Tuning & Parameter-Efficient Methods