Retrieval systems · measured, not marketed

RAG Systems
Evaluation Bench

A CPU-only evaluation of lexical, dense, and hybrid retrieval with transparent RAG failure analysis.

CPU-only$0 API costReproducible BM25MiniLMRRF

01 / Dataset

Benchmark overview

Original synthetic technical content. Small and diagnostic—not an industry benchmark or a claim of statistical significance.

02 / Retrieval

Same evidence. Three strategies.

RetrieverRecall@1Recall@3Recall@5MRR@5RuntimePeak RSS

Highlights mark the highest observed value per quality column. Runtime and memory were observed once on the CPU-only development machine.

03 / Diagnostics

Category breakdown

Each slice contains only a few examples. Use it to inspect behavior, not infer statistical significance.

Hybrid retrieval

Rankings speak a common language.

BM25 and cosine scores live on incompatible scales. Reciprocal Rank Fusion combines positions instead—without training or score calibration.

BM25 ranking+Dense rankingRRF RRF(d) = Σ 1 / (60 + rankᵣ(d))

04 / Inspection

Failure explorer

05 / Generation

Retrieval success is not answer success.

Retrieval failure

The labeled evidence never reached the model.

Generation failure

Evidence was retrieved, but the answer missed the reference under this benchmark's exact-match rule.

Reference run: BM25 + FLAN-T5-small. Semantic similarity measures meaning overlap, not faithfulness.

06 / System

Inspectable from source to failure.

Documents→Chunks→BM25 / MiniLM / RRF→Retrieval metrics→Context→FLAN-T5→Answer metrics→Failure type

Methodology

Deliberately modest.

Twenty-four curated questions test identifiers, acronyms, paraphrases, distractors, semantic wording, and controls. All systems receive identical chunks and labels.