Privacy

Accept optional first-party analytics or decline. Functional journey and sound preferences stay on this device.

Read the privacy notice
WorkshopSeptember 26, 2026

Building an AI system you can actually measure

A minimum viable evaluation setup for RAG and assistants, and what to do when the numbers say your design decision was wrong.

I benchmarked my own retrieval default and it lost. We build the same kit: a hand-labeled question set, an ablation across configurations, answer-level evals with the model on and off, and latency and cost from traces.

Conferences, meetups and internal engineering days are welcome. Invite me through the contact page.