Dev.to
7/11/2026

The original title is 12 words: "What happens when you ask 8 AI models the same buying question every month"
Original: What happens when you ask 8 AI models the same buying question every month
Short summary
The author built a monthly harness that asks 8 AI models to pick the single best tool across 16 B2B software categories. Models never unanimously agreed on any category, and individual models contradicted their own prior picks ~74% of the time across sessions. All data and prompts are open-sourced under CC-BY, with implementation notes for building your own version.
- •8 models tested across 16 B2B categories monthly — zero unanimous picks ever
- •Same model contradicts its own top pick ~74% of the time in fresh sessions
- •Full dataset and prompts open-sourced under CC-BY for reproducibility
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



