Evaluation lab
Benchmarks, in the open.
A small, repeatable view of the latest OurToken model canaries. Scores are the percentage of fixed test cases answered correctly.
Current canaryMMLU-Pro + IFEval2 models with a completed run
Latest results
Model scores
GPT 5.6 LunaOpenai ·
openai/gpt-5.6-luna30/ 1003/10cases passedAug 4, 2026, 9:02 AMCompleted with errorsDeepSeek V4 FlashDeepseek · deepseek/deepseek-v4-flash70/ 1007/10cases passedAug 4, 2026, 9:01 AMLatest runResults combine the latest completed MMLU-Pro and IFEval starter canaries available for each model. This is directional evidence, not an official full-suite ranking.