Key takeaways
- Benchmark saturation has pushed vendors toward task-completion metrics.
- Buyers increasingly ask for completion rate and intervention count, not benchmark scores.
- Tool-use reliability and recovery from errors matter more than reasoning depth.
Benchmarks stopped differentiating
With frontier models clustered within a few points of each other on public benchmarks, those numbers no longer help buyers choose. Vendors have responded by publishing task-completion data from long-running agent runs instead.
The useful question for a buying committee is now simple: out of 100 realistic tasks, how many finish correctly without a human stepping in?
What to ask a vendor
Ask for completion rate on tasks resembling yours, the average number of human interventions per run, and the recovery behaviour when a tool call fails.
Run a two-week pilot with your own tasks. Vendor-supplied demos are optimised for the happy path and rarely predict production behaviour.
References
- [1] Agent evaluation methodology notes — sortblogging research
- [2] Vendor reliability disclosures, Q2 2026 — industry filings
Frequently asked questions
Are agents production-ready?
For bounded, reversible tasks with review steps, yes. For unsupervised high-stakes actions, not yet.