A living atlas of AI evaluation
A model can be 99.9% accurate and catch zero fraud. Real trust is earned one question at a time — from framing the problem to keeping it honest long after it ships.
Question assumptions. Build intuition. Create impact.
Stop asking what a metric means. Start asking where it belongs.
Every metric — accuracy, F1, AUROC, calibration — exists because someone asked a question the last one couldn't answer. Learn the questions, and the metrics finally have a home.
Read the full essay →Ten questions every model must survive
Not one of them is a metric. Metrics are just the tools we use to answer them. A landmark earns its place only if skipping it can sink a model that looks impressive.
What problem are we solving?
Can we trust our data?
What should the model learn?
Is it actually learning?
How well does it perform?
Will it work elsewhere?
Should humans rely on it?
Can we safely deploy it?
Is it still working?
How do we keep improving?
We start with Evaluation
Evaluation is where a model looks most convincing and can mislead you most completely — and it's where all of my writing lives so far. We'll explore it one landmark at a time: performance, ranking, calibration, decision thresholds, and evaluating AI with AI.
Browse the essays on Medium →Three essays to begin with
Traveling with an office cat. Every good expedition needs one. He turns up across the atlas — pointing at charts, squinting at confidence scores, occasionally unconvinced. If a figure looks too clean, he's probably already suspicious.
Follow the journey
New landmarks land most weekends — a deep dive into one question at a time. Subscribe or follow wherever you read.