AI governance gets difficult when someone asks how you'll check whether a system works as it should. The principles are easy to agree with. You still need evidence. This book is built on 8 questions, asked in the order a system meets them across its lifecycle. Why did the model produce this result? Does the system treat groups differently? Is the output good enough for the task? How does it fail when someone tries to misuse it? What checks can prevent harm while it's running? What is it doing once people use it? What changes when it can take actions? And what evidence does someone need to approve it? One chapter per question: what it means, what to look for, and what good looks like - with hand-drawn figures throughout, the key lines bolded, and a short "where to start" at each chapter's close naming the open-source tools (20+ across the book, from LIME and Fairlearn to garak, tau-bench, and AI Verify). No code. You can read it in one sitting and walk into the next review or vendor meeting knowing what to ask. It is written by the person who led quantitative model risk inspections and AI risk supervision at the Monetary Authority of Singapore, and wrote Singapore's AI risk management guidelines for the financial sector. The questions are the ones he uses. And a new tool won't settle what fair means in your setting, or how many errors you're prepared to accept - someone still has to make those decisions and explain them. This book is how you get there. All of it is deliberately boring. That is the point. About forty pages. A short primer is listed separately, if you want the questions first.
AI governance gets difficult when someone asks how you'll check whether a system works as it should. The principles are easy to agree with. You still need evidence. This book is built on 8 questions, asked in the order a system meets them across its lifecycle. Why did the model produce this result? Does the system treat groups differently? Is the output good enough for the task? How does it fail when someone tries to misuse it? What checks can prevent harm while it's running? What is it doing once people use it? What changes when it can take actions? And what evidence does someone need to approve it? One chapter per question: what it means, what to look for, and what good looks like - with hand-drawn figures throughout, the key lines bolded, and a short "where to start" at each chapter's close naming the open-source tools (20+ across the book, from LIME and Fairlearn to garak, tau-bench, and AI Verify). No code. You can read it in one sitting and walk into the next review or vendor meeting knowing what to ask. It is written by the person who led quantitative model risk inspections and AI risk supervision at the Monetary Authority of Singapore, and wrote Singapore's AI risk management guidelines for the financial sector. The questions are the ones he uses. And a new tool won't settle what fair means in your setting, or how many errors you're prepared to accept - someone still has to make those decisions and explain them. This book is how you get there. All of it is deliberately boring. That is the point. About forty pages. A short primer is listed separately, if you want the questions first.