01 / Dataset
Benchmark overview
Original synthetic technical content. Small and diagnostic—not an industry benchmark or a claim of statistical significance.
02 / Retrieval
Same evidence. Three strategies.
| Retriever | Recall@1 | Recall@3 | Recall@5 | MRR@5 | Runtime | Peak RSS |
|---|
Highlights mark the highest observed value per quality column. Runtime and memory were observed once on the CPU-only development machine.
03 / Diagnostics
Category breakdown
Each slice contains only a few examples. Use it to inspect behavior, not infer statistical significance.
Hybrid retrieval
Rankings speak a common language.
BM25 and cosine scores live on incompatible scales. Reciprocal Rank Fusion combines positions instead—without training or score calibration.
RRF(d) = Σ 1 / (60 + rankᵣ(d))
04 / Inspection
Failure explorer
05 / Generation
Retrieval success is not answer success.
The labeled evidence never reached the model.
Evidence was retrieved, but the answer missed the reference under this benchmark's exact-match rule.
Reference run: BM25 + FLAN-T5-small. Semantic similarity measures meaning overlap, not faithfulness.
06 / System
Inspectable from source to failure.
Methodology
Deliberately modest.
Twenty-four curated questions test identifiers, acronyms, paraphrases, distractors, semantic wording, and controls. All systems receive identical chunks and labels.
- Small synthetic diagnostic benchmark
- No LLM judge or automatic faithfulness model
- No claim of production-scale performance
- Category groups are not statistically powered
- Public benchmark is distinct from the two-question integration fixture