Staging environment

What does a full eval suite look like in practice

Hosted by Madalina Turlea and Catalina Turlea

74 students

In this video

What you'll learn

Read a finished real eval suite backwards

Open a complete suite for a live AI feature and decompose top-line accuracy, cost and latency.

Decide which layer each failure belongs in

Some failures are rules check, some need an LLM judge, some stay human-annotated. Learn how to spot the difference.

See what a regression actually looks like

Compare iterations against the full dataset and watch what one extra prompt instruction quietly breaks elsewhere.

Why this topic matters

Eval suites get talked about constantly and shown almost never. Teams build one or two scorers, run them against the cases they already know about, and call it done. We'll read a complete eval suite of a live AI feature and show the behind the scenes: quality score, the criteria under it, then the dataset, then the showing the full iteration history. You leave with a picture of what finished looks like and a list of what's missing from yours.

You'll learn from

Madalina Turlea

Co-founder & CPO @Lovelaice, 10+ years in Product in FinTech

I'm co-founder of Lovelaice and a product leader with 10+ years building products across fintech, payments, and compliance. I hold a CFA charter and have led AI product development in highly regulated environments, where AI failures aren't just embarrassing, they're liabilities.

I've watched smart product teams make the same mistakes: choosing models based on benchmarks that don't reflect their use case, writing prompts that work in demos but fail in production, and leaving domain experts and PMs out of the AI iteration loop.

Through these failures (my own included), I developed a systematic approach to AI experimentation that puts product and domain expertise at the center. I teach what I've learned building Lovelaice: how to test, evaluate, and iterate on AI, before it reaches your users.

Catalina Turlea

Founder & CEO @Lovelaice | Co-founder & CTO @nilo | 14 years in tech

I bring over 14 years of software development expertise and a decade of startup experience to help teams build AI products that actually work. After founding my first company in 2019, I ran a consultancy in 2025, specializing in helping startups build MVPs, solve complex technical challenges, and integrate AI effectively.

I've seen firsthand how AI projects fail due to lack of experimentation and the right tools. Teams treat AI like traditional software and struggle with inconsistent results. That's why I co-created Lovelaice, a platform designed for non-technical professionals to experiment with AI agents systematically.

See all products from Madalina

Go deeper with a course

AI Evals for Product Managers: From Vibe Check to Proof
Madalina Turlea and Catalina Turlea
View syllabus