Privacy

Accept optional first-party analytics or decline. Functional journey and sound preferences stay on this device.

Read the privacy notice

Open to talks and workshops

RepoOctober 10, 2026

retrieval-eval

| config | R@1 | MRR | |---|---|---| | keyword | 87% | 0.922 | | semantic | 90% | 0.925 | | hybrid (RRF) | 87% | 0.898 | | semantic + rerank | 90% | 0.927 | | hybrid + rerank (shipped default) | 84% | 0.882 | 31 hand-labelled questions. The write-up explains why that is not enough to crown a winner.

Small retrieval evaluation for RAG systems: recall@k, MRR, nDCG, latency, a disagreement grid and a CI regression gate. No dependencies, no LLM judge.

I shipped a hybrid retriever on by default without measuring it. This harness showed it was the weakest of five configurations.