MarkTechPost
7/19/2026

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Short summary
Perplexity AI has released WANDR, an open benchmark with 500 evidence-heavy tasks designed to evaluate whether research agents can discover many qualifying entities and back each with cited, re-verifiable evidence. Perplexity Search as Code currently leads the benchmark with 0.363 soft F1 and 0.133 hard F1 scores. The article is a brief announcement with minimal detail.
- •WANDR is an open benchmark with 500 evidence-heavy tasks for evaluating research agents
- •Tests ability to discover qualifying entities with cited, re-verifiable evidence
- •Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



