Back to feed
Dev.to
Dev.to
7/4/2026
My AI memory benchmark said 98.3%. The number was true — and worthless.

My AI memory benchmark said 98.3%. The number was true — and worthless.

Short summary

The author benchmarked their Bastra Recall memory system at 98.3% recall, then realized the benchmark was flawed—testing with exact trigger phrases instead of paraphrased queries. Redesigned testing with 6 personas and found embeddings lifted far-query recall from 63% to 80%, while trigger phrases added zero lift. Key lesson: honest benchmarks measure paraphrase survival, not string overlap.

  • Initial 98.3% recall was a tautology—testing with exact trigger phrases instead of real-world paraphrased queries
  • Redesigned benchmark using 6 personas and paraphrased queries; embeddings lifted far-query recall from 63% to 80%
  • Trigger-phrase feature added zero lift on paraphrased queries; separate retrieval gaps from ranking issues

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more