Experiment Tracking & MLOps (MLflow & Weights & Biases)
Learn how mlflow and wandb track experiment parameters, metrics, and models so that model development stays reproducible and comparable.
Introduction
By this point in the course you have trained models with scikit-learn, built neural networks, worked with NLP and computer vision libraries, and just now learned how to schedule pipelines with Airflow and Prefect. There is one problem none of those tools solve on their own: after your tenth, fiftieth, or two-hundredth training run, which one was actually the best, and what exact parameters produced it?
Experiment tracking libraries answer that question. They record every run's parameters, metrics, and resulting model in one place, so you can compare runs instead of relying on scattered notebook cells and memory.
- How MLflow tracks parameters, metrics, and models, locally or self-hosted.
- How Weights & Biases (wandb) provides the same tracking as a hosted service with rich dashboards.
- How this ties together the modeling work from earlier lessons in this course.
mlflow: Local & Self-Hosted Tracking
MLflow is an open-source experiment tracking tool that you can run entirely locally, with no account or external service required, or self-host on your own infrastructure. Wrapping a training run in a few mlflow calls logs everything about it — parameters, metrics, and the trained model itself — to a local folder or a tracking server you control.
pip install mlflow scikit-learnimport mlflowimport mlflow.sklearnfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import accuracy_score
with mlflow.start_run(): n_estimators = 150 max_depth = 8
mlflow.log_param("n_estimators", n_estimators) mlflow.log_param("max_depth", max_depth)
model = RandomForestClassifier(n_estimators=n_estimators, max_depth=max_depth) model.fit(X_train, y_train)
accuracy = accuracy_score(y_test, model.predict(X_test)) mlflow.log_metric("accuracy", accuracy)
mlflow.sklearn.log_model(model, "churn_model")mlflow uiClick Run to see what this code prints.
Because MLflow logs the exact parameters alongside the metric and the model artifact, you can look back at run_d4e5f6 weeks later, see it used n_estimators=150 and max_depth=8, and reload that exact saved model with mlflow.sklearn.load_model().
wandb: Hosted Tracking & Dashboards
Weights & Biases (installed as the wandb package) solves the same core problem as MLflow, but as a hosted service with a strong emphasis on rich, real-time dashboards — live-updating loss curves, comparison tables across dozens of runs, and easy sharing of results with a team via a link, without anyone needing to run their own tracking server.
pip install wandbimport wandbfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.metrics import accuracy_score
wandb.init(project="churn-prediction", config={ "n_estimators": 150, "max_depth": 8,})config = wandb.config
model = RandomForestClassifier( n_estimators=config.n_estimators, max_depth=config.max_depth,)model.fit(X_train, y_train)
accuracy = accuracy_score(y_test, model.predict(X_test))wandb.log({"accuracy": accuracy})
wandb.finish()Click Run to see what this code prints.
MLflow is free, open-source, and keeps your tracking data under your own control — a natural default for local work or self-hosted infrastructure. wandb trades that self-hosting for a polished hosted dashboard and easier team collaboration, with a free tier for individuals and paid tiers for teams.
Common Mistakes
- Only logging the final metric and skipping parameters, which makes a good run impossible to reproduce later.
- Forgetting to call wandb.finish() (or leaving an mlflow run open), which can leave a run marked as incomplete.
- Tracking experiments in a spreadsheet by hand instead of using either tool, which does not scale past a handful of runs.
- Logging secrets or raw customer data as parameters or artifacts, which then get stored (and potentially shared) inside the tracking system.
Best Practices
- Log every parameter that could affect the result, not just the ones you expect to matter.
- Log the trained model itself as an artifact, not just its metrics, so a winning run can be reloaded directly.
- Give runs and projects descriptive names so past experiments are easy to find months later.
- Pick MLflow when you want full control and no external dependency; pick wandb when dashboards and team sharing matter more than self-hosting.
Frequently Asked Questions
It is unusual but not impossible — most teams pick one as their primary tracking system to avoid duplicated effort and a confusing split source of truth.
It's optional for a single quick experiment, but becomes valuable the moment you start comparing more than a few runs — which happens faster than most people expect once you start tuning hyperparameters.
No — mlflow has logging integrations for most major frameworks, including PyTorch, TensorFlow/Keras, and XGBoost, not just scikit-learn.
Key Takeaways
- MLflow tracks parameters, metrics, and models locally or on infrastructure you control.
- wandb provides the same tracking as a hosted service with rich, shareable dashboards.
- Both turn a pile of scattered training runs into a comparable, reproducible history.
- Experiment tracking is the piece that ties together every modeling library covered earlier in this course.
Summary
MLflow and wandb close the loop on the modeling libraries from earlier in this course — they are how you keep track of which model, trained with which parameters, actually performed best. In the final lesson, we bring everything from this course together into one realistic project and dependency stack.
- You can log parameters, metrics, and models with MLflow.
- You can track and visualize runs with wandb.
- You are ready for the final, capstone lesson of this course.