Dev.to
7/4/2026

My AI memory benchmark said 98.3%. The number was true — and worthless.
Short summary
The author benchmarked their Bastra Recall memory system at 98.3% recall, then realized the benchmark was flawed—testing with exact trigger phrases instead of paraphrased queries. Redesigned testing with 6 personas and found embeddings lifted far-query recall from 63% to 80%, while trigger phrases added zero lift. Key lesson: honest benchmarks measure paraphrase survival, not string overlap.
- •Initial 98.3% recall was a tautology—testing with exact trigger phrases instead of real-world paraphrased queries
- •Redesigned benchmark using 6 personas and paraphrased queries; embeddings lifted far-query recall from 63% to 80%
- •Trigger-phrase feature added zero lift on paraphrased queries; separate retrieval gaps from ranking issues
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



