AI Evaluation
Evals, benchmarks, retrieval quality, and building an evaluation culture.
Source episodes
Where this knowledge comes from.
Ep 85 - Leverage Outruns Wisdom: Systems Leadership In The AI Era
AI is quietly rewriting the org chart, and it’s not because everyone suddenly works faster. The real shift is structural: teams are becoming blended systems of humans, AI agents, orchestration layers, evaluation pipelines, and continuous automation workflows. Tha
Ep 84 - The Philosophical Shift: As Intelligence Becomes Cheap, Evaluation Becomes Everything
AI can generate code, analysis, and recommendations faster than any team in history, but there’s a catch: verification doesn’t scale the same way. When intelligence becomes abundant, judgment becomes scarce, and that scarcity reshapes what “good engineering” and
Ep 83 - Up the Stack: The Five Layers Of The Future Software Engineer
AI can write code faster than any team on earth, so why does it still feel like shipping software is hard? The uncomfortable answer is that speed is not the same as progress, and generation is not the same as judgment. We challenge the tired question “Will AI rep
Ep 67 - RAG Done Right: Measure The Evidence Or Drift Into Error
What happens when a brilliant-sounding AI gives the wrong answer with total confidence? We dig into the quiet culprit behind so many “LLM failures”: retrieval. Rather than judging how smart a model sounds, we walk through how to judge whether it looked at the rig
Ep 64 - Intelligence, Accountability, And You: From AI Slop to Sound Judgement
The pace of AI can feel exhilarating until a polished report collapses under scrutiny and your team spends hours repairing “work slop.” We’re seeing a quiet shift across organizations: as intelligence becomes ambient, leadership’s edge moves from gathering inform
