
$59
PDF · 128 pages · 1.03 MB
Secure checkout. Your download link appears the moment payment is confirmed.
Ravens AI Academy Series
Evaluating Intelligence: A Practitioner's Guide to AI Evaluation
A Practical Framework for Trustworthy AI — Metrics, Fairness, Safety, Red-Teaming, Governance and Production Monitoring
Organisations buying and building AI have no shortage of enthusiasm, and no shortage of vendors promising state-of-the-art performance. What they lack is a shared, practical language for proving whether a system actually works — and keeps working — after it touches real people.
Evaluating Intelligence closes that gap. It moves from foundational concepts through system-specific evaluation methodology, cross-cutting technical and ethical concerns, governance and monitoring, and finally into implementation guidance, real-world case studies and nineteen appendices of templates, rubrics and checklists you can put to work immediately.
Inside the book
- Part I — Foundations: the stakes of evaluation, what "good" means, and the evaluation lifecycle
- Part II — Evaluation by system type: classical ML, generative AI and LLMs, computer vision, forecasting, autonomous agents
- Part III — Cross-cutting concerns: technical metrics, fairness and bias, safety and red-teaming, explainability, security, human evaluation, benchmarking
- Part IV — Governance and production: GRC alignment, continuous monitoring and MLOps
- Part V — Implementation: building an evaluation programme, case studies, industry playbooks, team roles and the road ahead
- Appendices A–S: glossary, metrics quick reference, report and model-card templates, risk-tiering rubric, vendor questionnaire, maturity self-assessment and more