AI Evals in Practice: Testing, reliability, and quality for LLM systems in production

AI Evals in Practice: Testing, reliability, and quality for LLM

ISBN

Publisher

Imprint

Year Published

Print Length

Format

SKU

9781808817250
Packt Publishing
2026
150
26842

Original price was: ₨11,175.00.Current price is: ₨1,770.00.

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines Key Features Build a complete Python eval harness from datasets and scorers to CI and dashboards Evaluate RAG,…

Description

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines Key Features Build a complete Python eval harness from datasets and scorers to CI and dashboards Evaluate RAG, agents, and prompts with calibrated metrics and human feedback Move from offline testing to production evals, guardrails, and reliability workflows Book Description
LLM applications can look healthy in dashboards while their answers quietly become less accurate, less useful, or less reliable. AI Evals in Practice gives developers and AI engineers a systematic way to measure quality, catch regressions, and make evidence-based improvements before and after deployment.
You will build evalkit, a complete Python evaluation harness, while learning the core components of an eval: datasets, scorers, runners, and golden test sets. You will create deterministic scorers and LLM-as-judge evaluations, then calibrate judges and mitigate common biases. The book applies these foundations to prompt regression testing and CI, RAG retrieval and generation metrics, agent trajectories and tool calls, and human annotation workflows. You will then extend evaluation into production with online sampling, guardrails, and cost- and latency-aware pipelines, while comparing tools such as DeepEval, promptfoo, Langfuse, and Braintrust. A case study and final project bring the pieces together into an eval-driven development workflow with CI and a dashboard.
By the end, you will be able to design evaluation pipelines that help you ship LLM systems with measurable, repeatable quality. What you will learn Design eval datasets, scorers, runners, and golden test sets Build deterministic metrics for repeatable quality checks Calibrate LLM judges and reduce common evaluation bias Add prompt regression tests to continuous integration Measure retrieval and generation quality in RAG systems Evaluate agent trajectories, tool calls, and multi-turn behavior Run human annotation workflows and production online evals Control evaluation cost and latency without losing signal Who this book is for
This book is for AI engineers, LLM application developers, machine learning engineers, platform engineers, and technical leads who build or operate production systems using large language models. It is especially useful for teams working with prompts, RAG pipelines, agents, or AI features that need measurable quality and regression protection. Readers should be comfortable with Python and familiar with building or integrating LLM applications. Table of Contents Why Evals Are the Missing Discipline Anatomy of an Eval: Datasets, Scorers, and Runners Building Golden Datasets Deterministic Scorers LLM-as-Judge: Design, Calibration, and Bias Prompt Regression Testing and CI Integration Evaluating RAG: Retrieval and Generation Metrics Evaluating Agents: Trajectories, Tool Calls, and Multi-Turn Human-in-the-Loop Annotation and Labeling Ops Production: Online Evals, Sampling, and Guardrails Cost and Latency of Eval Pipelines Tooling Landscape: DeepEval, promptfoo, Langfuse, Braintrust Eval-Driven Development Workflow Case Study: Taking a Flaky Agent to 99% Reliability Final Project: Complete Evalkit with CI and Dashboard

Praise and Reviews

Not available

About the Author

AI Evals in Practice: Testing, reliability, and quality for LLM

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines Key Features Build a complete Python eval harness from datasets and scorers to CI and dashboards Evaluate RAG,…

Description

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines Key Features Build a complete Python eval harness from datasets and scorers to CI and dashboards Evaluate RAG, agents, and prompts with calibrated metrics and human feedback Move from offline testing to production evals, guardrails, and reliability workflows Book Description LLM applications can look healthy in dashboards while their answers quietly become less accurate, less useful, or less reliable. AI Evals in Practice gives developers and AI engineers a systematic way to measure quality, catch regressions, and make evidence-based improvements before and after deployment. You will build evalkit, a complete Python evaluation harness, while learning the core components of an eval: datasets, scorers, runners, and golden test sets. You will create deterministic scorers and LLM-as-judge evaluations, then calibrate judges and mitigate common biases. The book applies these foundations to prompt regression testing and CI, RAG retrieval and generation metrics, agent trajectories and tool calls, and human annotation workflows. You will then extend evaluation into production with online sampling, guardrails, and cost- and latency-aware pipelines, while comparing tools such as DeepEval, promptfoo, Langfuse, and Braintrust. A case study and final project bring the pieces together into an eval-driven development workflow with CI and a dashboard. By the end, you will be able to design evaluation pipelines that help you ship LLM systems with measurable, repeatable quality. What you will learn Design eval datasets, scorers, runners, and golden test sets Build deterministic metrics for repeatable quality checks Calibrate LLM judges and reduce common evaluation bias Add prompt regression tests to continuous integration Measure retrieval and generation quality in RAG systems Evaluate agent trajectories, tool calls, and multi-turn behavior Run human annotation workflows and production online evals Control evaluation cost and latency without losing signal Who this book is for This book is for AI engineers, LLM application developers, machine learning engineers, platform engineers, and technical leads who build or operate production systems using large language models. It is especially useful for teams working with prompts, RAG pipelines, agents, or AI features that need measurable quality and regression protection. Readers should be comfortable with Python and familiar with building or integrating LLM applications. Table of Contents Why Evals Are the Missing Discipline Anatomy of an Eval: Datasets, Scorers, and Runners Building Golden Datasets Deterministic Scorers LLM-as-Judge: Design, Calibration, and Bias Prompt Regression Testing and CI Integration Evaluating RAG: Retrieval and Generation Metrics Evaluating Agents: Trajectories, Tool Calls, and Multi-Turn Human-in-the-Loop Annotation and Labeling Ops Production: Online Evals, Sampling, and Guardrails Cost and Latency of Eval Pipelines Tooling Landscape: DeepEval, promptfoo, Langfuse, Braintrust Eval-Driven Development Workflow Case Study: Taking a Flaky Agent to 99% Reliability Final Project: Complete Evalkit with CI and Dashboard

Praise and Reviews

Not available

About the Author

Shopping Cart
Your cart is currently empty!.

You may check out all the available products and buy some in the shop.

Continue Shopping
Add Order Note
Estimate Shipping
AI Evals in Practice: Testing, reliability, and quality for LLM systems in production