The same questions, several free models, and a judge model that scores the answers without knowing who wrote them. Every run gets a link you can share.
Scores come from one model reading answers labelled A to D, shuffled, told that a wrong fact caps the score at 4. It is an opinion, not a benchmark: run it twice to see how steady it is.