Kolibrí
📐Paid course

Evaluating LLM Systems

How to know whether your prompt change improved anything or broke something else: datasets, LLM-as-judge and its biases, and regression suites.

📐

Evaluating LLM Systems

🎓 3 levels📚 12 lessons

1. Why You Need Evals

  • The problem: you do not know whether your change improved anything
  • Your evaluation set: where it comes from and how to build it
  • What can actually be measured, and what cannot
  • The small-numbers mistake

2. LLM as a Judge

  • LLM as a judge: how it works and why it is so widely used
  • The judge's biases, and what to do about each
  • Calibrating the judge against humans
  • Rubrics that produce agreement

3. Evals in Your Workflow

  • The regression suite: making sure a fix does not break what worked
  • Offline and online: two different questions
  • Traces: without observability there is no diagnosis
  • Your minimum eval system, running this week

$10one-time payment

Interested in more courses?

Kolibrí Pro gives you the entire catalog, including new courses as they launch, from $19/mo.

See all plans
See all courses
Evaluating LLM Systems — Online course with certificate