Evaluation lab

Benchmarks, in the open.

A small, repeatable view of the latest OurToken model canaries. Scores are the percentage of fixed test cases answered correctly.

Current canaryMMLU-Pro + IFEval2 models with a completed run

Latest results

Model scores

0–100 · click a model to view its catalog page

Results combine the latest completed MMLU-Pro and IFEval starter canaries available for each model. This is directional evidence, not an official full-suite ranking.