Back to feed
Dev.to
Dev.to
7/12/2026
Benchmarking Markdown Knowledge Graphs as Agent Memory: Lessons from LOCOMO

Benchmarking Markdown Knowledge Graphs as Agent Memory: Lessons from LOCOMO

Original: The benchmark that built the tools

Short summary

IWE's team benchmarked a markdown knowledge graph as agent memory against the LOCOMO dataset, the standard in published memory-system research. The benchmark repeatedly proved their initial hypothesis wrong, forcing rebuilds of their search engine and editing primitives. They ended with a $4.50 curation model whose store reads back at 96% of a hand-built ceiling in a single retrieval call, while grep over raw transcripts still wins on raw accuracy.

  • Benchmarked markdown knowledge graph as agent memory using the LOCOMO dataset
  • Benchmark failures drove rebuilds of search, editing primitives, and design rules
  • Final result: $4.50 curation model achieves 96% of hand-built ceiling; grep baseline still unbeaten on accuracy alone

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more