Evals and Aliens – How model testing is not a binary affair

Nov 17, 2025

The Confusion Matrix

00:00 / 01:05:14

Pete and Alex examine AI model evaluation methodologies, comparing traditional machine learning metrics with the qualitative assessment challenges of large language models. They discuss the collaborative requirements between technical and business teams to establish evaluation criteria for generative AI systems, highlighting the subjective nature of testing conversational outputs versus binary classification tasks. With the help of Ada, they also establish, once and for all, that Alien is a horror movie rather than sci-fi.

The Talk Python To Me podcast referenced.

Vanishing Gradients podcast.