Dev.to
7/12/2026

Benchmarking Markdown Knowledge Graphs as Agent Memory: Lessons from LOCOMO
Original: The benchmark that built the tools
Short summary
IWE's team benchmarked a markdown knowledge graph as agent memory against the LOCOMO dataset, the standard in published memory-system research. The benchmark repeatedly proved their initial hypothesis wrong, forcing rebuilds of their search engine and editing primitives. They ended with a $4.50 curation model whose store reads back at 96% of a hand-built ceiling in a single retrieval call, while grep over raw transcripts still wins on raw accuracy.
- •Benchmarked markdown knowledge graph as agent memory using the LOCOMO dataset
- •Benchmark failures drove rebuilds of search, editing primitives, and design rules
- •Final result: $4.50 curation model achieves 96% of hand-built ceiling; grep baseline still unbeaten on accuracy alone
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



