Introducing Sontairo Benchmarks
Benchmarks · AI Models · Product Updates · Releases
Choosing an AI model shouldn’t feel like guesswork. Sontairo Benchmarks gives you a clearer view of which model performs best overall and for the work you care about.
Compare models by the work you actually do
Choose Overall for a general-purpose recommendation, or explore Planning, Coding, Design, Security, Research, Data analysis, Business, and Tool use. Each category covers multiple competencies, so one impressive answer cannot define a ranking.
You can switch between Best quality and Best value. Value compares strict quality against observed cost while keeping the quality score visible, so a cheap but weak answer cannot hide behind price alone.
Scores are deliberately stricter
High scores are now harder to earn. A raw rubric result is translated through a stricter customer-facing curve, and rankings appear only after the evidence clears gates for sample size, competency coverage, confidence, safety, and observed cost.

The new calibration makes 90+ exceptional: a raw 90 displays around 79, and visible scores are capped below 100.
Your feedback matters—without letting one vote swing a ranking
Thumbs up and thumbs down are the strongest real-use signal. They contribute only after enough evidence has accumulated across conversations, so one reaction cannot move a model to the top. Controlled testing remains the anchor when both evidence sources are available.
Efficiency is part of value
Benchmarks should measure more than headline quality. Sontairo also tracks whether models use context and tools efficiently. Recent improvements reduce repeated setup, load specialized tools more selectively, and avoid unnecessary model calls on focused work.

In one targeted workflow benchmark, the repeated-input cost index fell from 100 to 57 while the same evaluated outputs averaged 99.2 on the quality rubric.
These results are deliberately scoped. They describe the tested workflows—not every possible task—and Sontairo withholds recommendations when the evidence is not yet trustworthy.