— Three new models
I added Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna. They scored 75.6%, 55.5%, and 21.0%, respectively, on the same 119 questions with the same run settings. All three returned an answer to every question.
Opus 5.5 moves into second place behind GPT-6 Astra. The table, chart, downloads, and example answer pages now include all 13 models.

