Back to feed
MarkTechPost
MarkTechPost
7/19/2026
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

Short summary

Perplexity AI has released WANDR, an open benchmark with 500 evidence-heavy tasks designed to evaluate whether research agents can discover many qualifying entities and back each with cited, re-verifiable evidence. Perplexity Search as Code currently leads the benchmark with 0.363 soft F1 and 0.133 hard F1 scores. The article is a brief announcement with minimal detail.

  • WANDR is an open benchmark with 500 evidence-heavy tasks for evaluating research agents
  • Tests ability to discover qualifying entities with cited, re-verifiable evidence
  • Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more