All topics
Topic hub

AI Evaluation

Evals, benchmarks, retrieval quality, and building an evaluation culture.

Source episodes

Where this knowledge comes from.

Ep 85 - Leverage Outruns Wisdom: Systems Leadership In The AI Era

AI is quietly rewriting the org chart, and it’s not because everyone suddenly works faster. The real shift is structural: teams are becoming blended systems of humans, AI agents, orchestration layers, evaluation pipelines, and continuous automation workflows. Tha

Ep 84 - The Philosophical Shift: As Intelligence Becomes Cheap, Evaluation Becomes Everything

AI can generate code, analysis, and recommendations faster than any team in history, but there’s a catch: verification doesn’t scale the same way. When intelligence becomes abundant, judgment becomes scarce, and that scarcity reshapes what “good engineering” and

Ep 83 - Up the Stack: The Five Layers Of The Future Software Engineer

AI can write code faster than any team on earth, so why does it still feel like shipping software is hard? The uncomfortable answer is that speed is not the same as progress, and generation is not the same as judgment. We challenge the tired question “Will AI rep

Ep 67 - RAG Done Right: Measure The Evidence Or Drift Into Error

What happens when a brilliant-sounding AI gives the wrong answer with total confidence? We dig into the quiet culprit behind so many “LLM failures”: retrieval. Rather than judging how smart a model sounds, we walk through how to judge whether it looked at the rig

Ep 64 - Intelligence, Accountability, And You: From AI Slop to Sound Judgement

The pace of AI can feel exhilarating until a polished report collapses under scrutiny and your team spends hours repairing “work slop.” We’re seeing a quiet shift across organizations: as intelligence becomes ambient, leadership’s edge moves from gathering inform