Users can build test suites from their own data and use cases and compare models by quality, cost and speed per task. It targets the known weakness of public leaderboards.
Artificial Analysis has so far run public comparisons of language models. With Optima, users can now create their own test suites from internal data and concrete use cases and measure several models against each other by output quality, cost and response time per task.
The motivation is familiar: public leaderboards test general capability, not a company's specific task. A model at the top of a leaderboard can perform noticeably worse than a cheaper one when processing your own service requests, quotation texts or delivery notes.
That shifts the basis for decisions. Knowing your cost per completed task lets you deploy smaller models exactly where they suffice and reserve expensive models for the cases that need them. Without such measurement, vendor choice remains guesswork.
What this means for decision-makers
- Build a test suite from real, anonymised cases for your two or three most important AI use cases.
- Measure cost per completed task rather than price per million tokens – only that figure is comparable to your budget.
- Repeat the measurement after every vendor model change; results shift without any action on your side.
This story was produced automatically from the source named above and checked by software before publication. The image is symbolic and shows neither the event nor a real person. How this paper is made
